That was a terrible ruling. It should be about who got the camera there and set it up, not who pushed the button. With automated recording (dashcams) there isn't even "a button."
>It should be about who got the camera there and set it up, not who pushed the button. With automated recording (dashcams), there isn't even "a button".
I'm actually having a bit of trouble thinking of what sufficient societal good there is/would be in granting copyright on raw dashcam or security camera or the like footage? None of those purely mechanical automated systems need a subsidy or encouragement to generate more. Certainly someone can use that sort of thing in the creation of a copyrighted work but what would be the issue with the underlying material in that case being unprotected?
I am in favor of copyright reform so eventually it becomes something like two years automatic with a one time two year extension possible an I agree with you because of one key word — raw. Once the footage is used in a package, be it a movie, a news report, or a music video, that package can be copyrighted. The problem is there is no way the owner of a dashcam can prohibit commercial use of a video they own if it cannot be copyrighted in the US. In a way, we have conflated copyrights with other so called intellectual property (details). Maybe the solution is to strengthen privacy rights somehow? I'm not sure.
Thinking more about monkeys, I think that case was made relatively easy because there was a normal living recipient for copyright, but that wasn't eligible for such. I'm fairly certain that dashcam footage, which is 100% automated, as you mention, is copyrighted by default. It's a fairly common source of footage of novel events like a meteorite, rocket impact, or whatever else. And that footage is licensed to media companies.
In any case, any country that decides neural network generated media isn't copyrighted then faces a major, and probably impossible, problem of trying to prove it wasn't made by a human.
Theoretically, in the US dashcam, CCTV, and similar automated recordings are not copyrightable. In the US you can only copyright creative expression. However, the question hasn't been answered in court, yet, AFAIU.
That doesn't stop people from treating automated recordings as if they were copyrighted, though. Especially for media companies, licensing is standard operating procedure. Even if a media company's lawyers believe something isn't actually copyrightable, if the licensor and licensee believe it is, and especially if distribution outlets (e.g. YouTube) treat it as such, then it all works out.
In some other countries automated footage is copyrightable, AFAIU, but it varies.
> In any case, any country that decides neural network generated media isn't copyrighted then faces a major, and probably impossible, problem of trying to prove it wasn't made by a human.
In the US the initial burden is on the person claiming a copyright violation. In court that is met by simply showing a certificate from the Copyright Office. In fact, it's actually a requirement--you can't sue in court to enforce a copyright claim with a certificate. But in getting the certificate you would be making a declaration that it was made by a human.
After this ruling, I had an idea for a photo where I set up the entire set, camera, etc, but the photo entailed the model clicking the shutter while we were both in the shot. Just in case, I had the model write a quick note ceding the rights.
It's fine, I guess. How does it work in cinema? A director who is the creator of the project must have to get rights from every camera/mic operator.
I think the issue you then run into is imagine another human came, picked up the camera, and used it to take a picture of whatever. It'd be fairly self evident that that the photographer would own the copyright to that work, but in your take - it'd be the camera owner.
So you're more of trying to create a special rule where if the normal recipient of a copyright would be invalid, then it slides to the 'nearest' most appropriate individual, but that seems extremely fragile and difficult to define.
Dashcams and security footage definitely don’t have copyright protection. Are you seriously thinking that the construction worker who set up the camera would then have copyright over everything it records?
Really? Regardless of the debate on what should and shouldn't be copyrighted, in the current system there is no way to set up a continuous camera with intention to create something and not be afforded the same rights as if I took a snapshot?
As software, creative works, science, etc become more and more contributed via AI does that mean all future works will not be copyrighted or patented? Under our current copyright culture and laws obviously not but that does open the question to how much does a human have to contribute and what evidence is required to show that a human contributed enough so that it can be copyrighted and patented. Some time in the future (may be many years) science may become to complicated for humans to understand fully with AI systems researching themselves. Current patent laws in many places including US says inventions created entirely by artificial intelligence cannot be patented. This future may be coming. What will the new copyright and patent laws look like in the future? Do we need copyrights and patents anymore?
It means as long as you can find a few employees who testify that they made "creative input" during the process and nobody can disprove that, you're fine.
It's like things that are already in public domain. Even if you make a coloring book out of paintings in public domain, it doesn't necessarily mean others can just print your book as-is.
> It means as long as you can find a few employees who testify that they made "creative input" during the process and nobody can disprove that, you're fine.
Have you been involved in copyright or patent litigation?
It's usually easier than that, as one by default hold the copyright of their works. Typically a company doesn't need to provide the whole working process and files to the court to claim copyright over their published products / to accuse pirates for violation.
If you're on the defendant side of a copyright violation case, it's extremely hard to use "well the original author didn't really make it...* as a defense. (Patent cases are often defended with this argument though, as a patent grants far boarder protection than copyright and can be rejected on prior art. But still it's very different from "AI made this actually.")
> It's usually easier than that, as one by default hold the copyright of their works.
That's not what we're talking about.
This is a question of copyright eligibility in the first place. An excellent recent(ish) example is Oracle v Google wherein the former claimed copyright over the Java APIs. That particular litigation went on for years and ultimately went to the Supreme Court (though it ultimately ended on a fair use argument and the issue of API copyright eligibility was never settled).
We can probably assume similar high profile cases in the future that turn on quantifying the level of human contribution in code produced through the use of AI.
As for patents, it's literally that easy. Attacking authorship of a patent is an excellent way to invalidate it. It's why any company with a serious patent program requires staff to keep supporting notes for their work and confirm authorship in writing as part of the filing process.
You can blame the governments for all sort of inefficiencies but any operation that impacts taxes runs a tight ship in most developed nations, this include things like copyright law, census, and anything statistics.
If companies get slope with creative output to the point that "a few employees" can reproduce it in shadow markets, don't expect to get copyright protection without giving governments revenue and speech-control.
Perhaps but what about individual works done not via a company. Say you are a author and someone falsely detects a AI watermark in your book. Is the testimony you provide enough to override that falsely detected AI watermark? My feeling is that there will be an industry that arises that proves you contributed to the works you want copyrighted.
My understanding is that you get back a percentage from SynthID regarding how confident it thinks the text is generated by AI or not. The question is what percentage would be considered undeniable that it is AI generated or require more than just your testimony to prove that you actually did the work yourself. https://www.nature.com/articles/s41586-024-08025-4 (Google Deepmind) was using a "Watermark detectability is measured using the true-positive rate (TPR) when the false-positive rate (FPR) is set to 1%" for instance in their research paper. This 1% was set arbitrary just for the research paper but what percentage should it be set to?
I think you need to provide independent sources for the "SynthID has practically zero false positives". Except a paper written by Google authors about SynthID for images I have not found any study about zero false positives.
"Content that is entirely generated by artificial intelligence is not protected by copyright."
If that is faithful reading of the law, that makes sense. I know a number of people who use AI, but none of them (that are making anything actually useful) have the output "entirely generated" (aside from some POC tests that never see the light of day).
I have a hard time believing anything of value, anything worth copyrighting, could be entirely generated by AI.
The current trend (since months now) is to have everything done through agentic loops. Meaning humans are only here to provide the initial prompt and do a few back and forth during the implementation.
> I have a hard time believing anything of value, anything worth copyrighting, could be entirely generated by AI.
Perhaps you're not stretching your imagination enough. What if an expert novelist used an AI like it were a fancy auto-completing dictation machine to write the next great American novel? AI may have "entirely generated" all the text, but what if they micromanaged the shit out it?
I can imagine the difference between someone who fires off a lazy 5 minute prompt, and someone who labors for months and months to get exactly the results they want.
It’s almost like laboring for months and months would have draft revisions, dozens of changes, and many other interactions they could prove as theirs defeating your entire argument.
This is less relevant for logos & creative works, but things that enable companies to do production will go back to being closely guarded and sharded secrets, which is what the patent system was trying to resolve (though we can certainly argue the success of it).
If companies fail to protect their investments in generating IP, they will stop investing in generating it.
And unless IP generation costs (all in, including the humans telling them what to generate) fall close to zero, it will be bad for the world if companies cannot recoup investments in generating new IP.
We would expect this to hit those industries relying on IP protections the most, e.g. pharma.
Take note of the qualifier entirely. If you're working with an agent steering it to produce the results you want, it would be an entirely different story.
Wheres the limit. Would a prompt "Create a system that does x" and no other feedback count? Or do need to do 10 prompts to steer your agent to the results you want. What about just 1 additional prompt. Somewhere there will be a line.
Legal systems are fuzzy. The limit will be what courtrooms decide. There will never be a ">75% of content" law, everything will be left to interpretation, laws are not source code.
> Neither mere prompting nor the selection between several AI suggestions is sufficient as a human creative contribution.
which reads like you actually have to contribute by modifying the output from the system, i.e. prompts/inputs to the system do not count as a human contribution.
I think that you might be taking "knowledge" too literally. As I read it, the parent's post was using "knowledge" poetically for the various works that AI companies stole.
You wrote your house's security code on a Post-It note. Somebody came in as a houseparty guest and located the note, and discretely copied it onto their hand to get access to the house during your next vacation.
Which verb would describe their knowledge of this security code?
The knowledge wasn't obtained illegally or illicitly, immorally perhaps.
Illegal entry/trespass is a crime regardless of the method used however.
In your analogy copying the code (online content) isn't necessarily illegal, however using that knowledge without compensation, permission or attribution may be a crime (depending on copyright law, judges rulings etc).
Seeding is often considered the actual point of infringement, aka sharing copyrighted material without the right licence.
I should note I suppose that I don't agree with companies using artists or developers hard work for their own financial gain without compensation, whatever form that may take.
The problem as usual is the availability of legal loopholes when you have a team of lawyers and near infinite money.
It really depends the country, downloading a Hollywood movie via torrent is itself illegal in a lot of countries. The only exception I’m aware of is Switzerland
Whatever it is, slapping "steal" on it because what was done is wrong, and stealing is wrong, is the same sloppy logic that the copyright lobby used in the 80s/90s.
So this means you can not use AI for the majority of open source projects since GPL, MIT, BSD, etc are all copyright declarations and they'd be being made for code which you can not copyright.
You can't enforce those licenses against AI-generated parts, because enforcement relies on the recipient having no other way to avoid copyright infringement.
However, there's no issue with including non-copyrightable code in otherwise copyrighted projects. There's already plenty of non-copyrightable code like auto-generated boilerplate.
>no issue with including non-copyrightable code in otherwise copyrighted projects
That is not how copyright/trademark/contract laws work, and isomorphic plagiarism is not a long-term business model. People also loved Napster at first too. Good luck =3
I find it hard to find correct terminology in this case. AI generated content is copy right wise in state of nothingness. It simply does not have copy right status like other material can have. As such maybe best I can formulate is that you can't enforce license violations against ai generated material as you do not have ownership.
So no license is enforceable with code written by AI.
No, GPL is a contaminating license contract, so "AI" slop means you are probably in GPL violation by including isomorphic plagiarized uncopyrightable code. =3
wasnt there a case a while back, where GPL or LGPL code made its way into MIT licensed software via LLM? And they were forced to remove the copyleft code. I dont remember the details though.
Usually what I have seen is someone writes a Perl library, Claude mostly ports it to Python with isomorphic plagiarism, and a bunch of Brogrammers proclaim "AI" magic is real.
The issue is most GPL license fall under contract law, and scraped code can't legally have assigned "copy" rights on an "AI" vector search compaction output.
Indeed, but people will continue to think vector search compaction similarity absolves folks of Trademark and Copyright liability.
As the dark specter of Disney Mickey Mouse looms over every LLM model involved in isomorphic and character plagiarism. Yes, even motion capture is considered a performance act in the guilds, so video reskinning an unlicensed performance act people make is also a liability.
It would sure save a lot of money if you don't get caught, so people are gonna try it for sure. =3
We know for a fact that Bun port to Rust was done entirely by Claude! I’m wondering if that could be used against Anthropic in some ways (not that I want to, just curious what would be the angle)
It is actually pretty easy when the area is very specific like 1 guys Perl library, and Claude ports it to Python for a group unaware of what happened.
All models know what Disney Mickey Mouse looks like too. =3
This post is glossing over most of the nuance in EU law.
The AI system must function merely as a tool or instrument (like a camera or Photoshop) guided by the human, rather than acting as the creator itself. The line may get a bit fuzzy case-by-case, but effectively the human must be the creative one, not the AI.
This is not unprecedented. Machine generated technical data, sensor outputs, automated surveilance photography, monkey selfies, purely algorithmic or generative music and such were already disqualified long before AI came along.
Interesting thought experiment is to consider an author who writes a novel in English, but wants to translate it into German.
They have a copyright on the original, and if they hire a human, the human would have a copyright on the translation (which would generally be licensed or transferred back to the author in some way).
If they use an AI for the translation, by the logic here, the translation wouldn't have its own independent copyright, but (based on other long established principles of copyright) it would still be a derived work of the original, so even if this decision holds it would not be legal to make unauthorised AI translations, pirate authorised AI translations, make further translations into other languages (or back to English), etc.
Which seems fairly reasonable! But consider:
If you start with, say, a 90,000 word novel, and ask for a translated novel, you (presumably) have sufficient rights to stop someone making unauthorised copies of the AI translated version.
If you start with a 300 word prompt, and ask for a logo, you (apparently) do not have sufficient rights to stop someone from using it without authorisation.
So some combination of the input (0.3k vs 90k) and the output (logo versus novel) crosses an inflection point between these two extremes, and I think it's interesting to wonder what the boundaries are. Like, in theory you could graph input size versus output complexity, and sketch a frontier between "the author's protected expression survives in the output" and "the author's protected expression does not survive in the output". And I don't have the slightest idea what I think a fair frontier would look like.
> Some combination of the input (0.3k vs 90k) and the output (logo versus novel) crosses an inflection point, and it's interesting to wonder what the boundaries are
It has always been like that though. Copyright is just that arbitrary. You really can't tell if Android violates Oracle's copyright over Java by reading law text.
It's supposed to be arbitrary. Copyright is not some sort of human right or divine commandment that creators have. It's a legal device to incentivize a certain type of behavior. Like tax cuts and subsidies, governments can dangle it in front of people whenever it wants.
Monkeys and LLMs (so far) either don't understand or don't need such incentives, so they don't get copyright.
I've come to understand that rules are not executable, or not executed. As programmers we understand if we write if(X) {Y} and X happens, Y will.
Laws aren't like that. Laws are just powerful people telling us what we want to do and what they want to do to us if we don't. The difference is subtle: the powerful person's power is not limited by the law. He only writes the law because exercising his power is annoying and he'd prefer that we complied in advance so he doesn't have to. The thing that actually executes is the powerful person's desires, not the law.
Input size and output complexity are not necessarily the only two dimensions involved. There’s no reason to think that such graph would show a continuous frontier.
As far as I'm following it, the court didn't suggest the prompt was not a protected work.
But what the court did say is that the logo wasn't (sufficiently) derived from the prompt; those 300 words weren't contained or encoded in the logo, so if you copied the logo without permission from the author of the prompt, this was fine, because you weren't in any meaningful sense copying that prompt, or any other protected work of the author. (Which in my view is fair enough, really!)
Another way of thinking about it is to consider:
1) The number of different prompts that would yield a very similar looking logo. In many cases, a one or two word difference would probably yield a very different logo. (Conversely, you could rewrite an entire chapter of the novel, but the other chapters would still be translated more-or-less unchanged. You'd need to rewrite every sentence to end up with an entirely new set of sentences in the output, and since copyright covers the arrangement of text, not the plot or ideas, that matters.)
2) How feasible it'd be to look at the logo and accurately guess the prompt. (Again considering the hypothetical novel, if you translated the novel back into English, entire passages would likely be word-for-word identical with the original.)
it's sort of non-controversial at the extreme end I think - where there is no human contribution, why should there be protection?
But it just begs the actual question of how much human contribution there needs to be:
- I wrote the prompt (not enough)
- I wrote many prompts and iteratively refined them using distinctly human skill (open question, but loosely seems still not enough, potentially in the EU but maybe in the US?)
- I made minor modifications post-generation (open question, probably enough)
- I made equal or more contribution to the final result (this better clearly have copyright protection or we are in real trouble)
> I made minor modifications post-generation (open question, probably enough)
Hmm I don't think that would be enough. I'd expect that would only make the modifications themselves copyrightable, but not the whole modified work including the AI parts.
Compare for example the case where the US copyright office ruled that, when assembling AI-generated images and human-written text into a comic book, only the human-made elements themselves (text, arrangement) get copyright protection, but not the images.
E.g. if you generate an AI photo and color grade it, i expect only the color grading would be protected (if that is even significant enough to be protectable), not the rest. And someone else could re-color-grade the same image without infringing your copyright.
Is it fair to say, then, that artists who have trained on the work of prior artists should also not be able to protect their works unless they have directly received consent and have compensated all the artists whose works they have viewed and potentially learned from over their lifetime?
Does this also extend to generated source code? I mean, this upend the business models of many A.I. companies like OpenAI and Anthropic. In addition, many ISVs would probably stop using A.I. if their products could lose their copyright.
And what about A.I. generated code which has been altered by a human? Is the derived work copyrightable?
And this is one of the reasons some people are up in arms about Anthropic's recent inclusion of watermarking against anything Claude models touch. It will even take human-generated code that it works on and apply watermarks to it which can be then used to prove it was AI processed and perhaps invalidate any copyright on even human-written code.
Im not sure it’s really a copyright issue here, and not patent or trademark related, or something else. Copyright is really only one aspect of IP laws and they all have their own legal nuances
I don't see an easy test here. Worse, I see the beginnings of a test that is technical and very likely to not match the future of how we will interact with these things. We have to start somewhere but I'm not sure 'Neither mere prompting nor the selection between several AI suggestions is sufficient as a human creative contribution.' is the right place to start. I think we need more examples of what does constitute 'human-centric' and work from there. I also don't think that any system that says 'you didn't do enough work so it isn't human-centric' works. Any system like that will require a reexamining of copyright in general. There is a lot of low work copyrighted material out there. Just because 'AI' didn't build it doesn't mean the same tests shouldn't apply.
Some examples of why I think this is really hard: Say I build a story generation system. I work hard on building an agent swarm of actors, critics, editors, researchers. I craft into the various agents concepts of story arcs, outlining techniques, character development. I build a huge well thought out process for how to agentically write an actually good story, so long as you give it a title. Heck, I even design and train my own custom LLM with original layer ideas and novel training techniques to use on this system. After all that I then take that final step and give it a title. Do I have no claim to that? I probably put more work and creativity into it than an author would have a book. What if I then gave it 500 titles? 5,000? Would my claim degrade the more titles I fed it? Is it a percentage of work question? What is the core concept here that defines 'human-centric'? What is the cut-off here?
Let's go even further. I don't prompt. I live in a world with unlimited context models. I have a conversation about the book I want it to write. During that process I reject some ideas and accept others. I didn't give it a 'system prompt' but essentially all I did was prompt it and select versions I liked. Is that not human centric? How about if I asked it for advice and it did some editing work on my story? Did that make it not human centric even though the starting text was mine? What if that starting text was 99% replaced with a version 10x as verbose. Defining based on how you interacted with the model (prompted and selected) just seems way to weak to be a clear test.
Actually, in Germany, copyright remains with the author and their heirs until seventy years after the author's death; it cannot be sold or given away, even if someone paid them doing it. Instead, there is "Verwertungsrecht" or "Nutzungsrecht" ("License", "Usage right"). In the USA, copyright can be sold completely to another legal entity.
Information wants to be free. We should simply dispose of copyright. With LLMs, waters that were already muddy are now a thick slurry. And it's only going to get worse.
It's an antiquated mechanism and is far more abused than it is actually used at this point.
Lots of wild guesses about mixed human-AI copyright in here. Last time I read the copyright office’s statements in this, their stance was basically: The human owns exactly what they contribute. The rest is public domain.
Yes, that is vague. I think the examples were like:
If you paint a symbol and use an AI filter over that to stylize it, you own the symbol aspect of the image but not the stylized final result.
You can own a book of AI images as a curated collection. But, not the individual images.
Seems about right. Copyright is supposed to literally prevent outright copying. The output from an LLM is not a creative work of the prompter. It's genuinely the opposite, I use them when I don't care about something but have to do it anyway for whatever reason. It makes more time for me to do the things I like working on.
It feels shitty that they have been trained on the life sums of all of our work and online presences with absolutely no credit given... But then again, I'm not sure I'd want to know what parts of the weights were from me and which weren't.
I feel like people will be upset by this take... But you don't get credit for the creativity of an idea alone. Every asshole has ideas. It's called a work of art for a reason, emphasis on WORK.
I’m trying to imaging what giving credit would look like… “Credit: Everyone who has made anything freely observable on the internet.” Would folks be less upset if that was included?
Frankly, I should have been putting something similar attached to any drawings or paintings I make physically. My entire life I’ve been training my brain on countless works of others. And, I’ve never asked one of them for permission or given anyone any credit.
Suppose that an artist trains an AI model purely on their own works, and then the AI generates something similar to that artist's work. I would say that the artist must be allowed to assert copyright over that. The artist has copyright over all the training data, and the output of the AI is a derived work of that.
What do you think should happen if an artist views a lot of arts over their lifetime and then themself generates something similar to another artist's work? I would argue that collectively humans have been standing on the backs of other humans all throughout history. An artist today does not simply come out of the womb knowing how to make great art - they learn it by observing other art, learning about art and practicing (not altogether dissimilar to model training) over many, many hours.
The set of possible and desirable non-copyrighted text/audio/visual states to render is effectively infinite. Laws don't prevent Scrabble clones; tropes are not protected.
Endless remixes of public domain content are an option as well.
If an iPhone with model weights in chip, a local Mac mini with similar but more powerful model on chip ends up capable of generating endless content copyrights won't provide a moat.
I feel like this question is typed wrong but the same reason someone wants copyright on human work to make money by prevent other people from stealing your work
Given the cost of generating images/text/etc. is much lower when using an AI system, as compared to a human, this seems appropriate. The purpose of intellectual property is to encourage people to invest that time into creating works that benefit others. Perhaps there’s an argument for a shorter copyright period to cover the cost of the compute spent, but certainly one wouldn’t expect AI intellectual property to require a similar level of incentive to produce.
Year ago, there was a painter, Thomas Kinkade, the "Painter of light".
He became a bit of a pop phenomenon, with stores dedicated to selling high quality prints of his works.
And one of the things, so I'm told, that could happen, was that you could buy one of his prints, but someone at the store, an artist, would dab some paint onto the print. Add some "light" to it.
Obviously this is a commercial endeavor, so I doubt the provenance of the rights holder was in doubt (such as through an employment agreement).
But it's, perhaps, an interesting case study about who owns what in a time of augmented and manipulated media.
I believe those are called “embellishments”, and they’re semi-common as a way to make a print “unique”, but still cost way less than the original the print was copied from. From the couple artists I talked to about them, their students do these at their studios, rather than the artist themselves.
Yes it does because there's is still direct control from a human. When working with LLMs, even interactive prompting does not provide that level of control.
This has already also been implied in the US. Courts upheld the Copyright Office's stance on human creation in the context of AI image generation. There's no reason to expect something that fundamental to copyright to be any different for other media, such as source code.
It's reasonable to expect this to hold for all Berne Convention countries.
All licenses are unenforceable if you don't hold the copyright, since you don't have a copyright to license. And attempting to do so would probably be perjury.
Makes it even worse. Basically it hugely depends. For example Finland allows letting go of commercial exploitation rights and even make binding statement on that. But still there always remains right for them to be informed if copies are being published.
Not to even mention moral rights on how such work can or can not be changed if it would be disparaging for example...
Funnily I think copyright protection for LLM model is 15 years. As to me it just looks like very fancy database and that is the protection for those...
>The creativity requirements may seem arbitrary but there’s a legal distinction between a sculpture and a standard brick.
Or more relevantly, a recipe find on recipe sites (with the author's entire backstory) vs a sequence of instructions. The latter is not copyrightable, even if there was some creativity that went into it (eg. word choice).
Judges are not at all stupid enough to fall for that. You would get laughed out of court for that. Judges are not computers, and they can not be tricked with these kinds of technicalities.
In fact, the defense of "I wrote the prompts that led to the code that the LLM wrote" would be a much better defense.
My bet would be to build anything from multiple independently generated pieces of AI code. Thus the software build from AI generated blocks by human would at least have protection on basis of the structure and work put into that.
Interesting. In addition to Anthropic's watermark use to prevent model collapse, we can definitively call something copyrighted or not copyrighted. This is a boon to everyone who consumes culture.
That doesn't actually work. Say you took a book and used an AI to translate it into another language. The translation wouldn't have an additional copyright, the way it would if a human had done the work, but the output would still be restricted by the original copyright. So the presence of the watermark does not tell you that the text is public domain.
I thought translations can be copyrighted separately. Their are translators for instance who translate very old texts from other countries/languages and then sell that book/translation using copyright to protect this business models? Not sure but does this mean if you as a human wrote a book in English and then used AI to translate to another language say French then that translated work would not be copyrighted. Not sure how this would work.
Translation, in and of itself, is viewed as a creative work. A translated work has two copyrights: the original, and the translation, with permission needed from all rightsholders to redistribute. A new translation of a public domain work (Emily Wilson’s The Odyssey) has one copyright holder, the translator. An existing work (the English translation of The Three Body Problem) has at least two: the original author and the translator.
However! Since AI work is noncopyrightable, the AI’s effort in translation is simply ignored. Claude’s The Odyssey would have zero rightsholders, and remain public domain. ChatGPT’s translation of The Three Body Problem would still be under Liu Cixin’s copyright.
Lmao. No. Content that is wholesale generated by AI is also not subject to copyright in the US. It relies on the honor system and you can always modify the material juuuust enough that you can claim the copyright. In practical terms, this changes nothing.
It’s been established in the DC circuit, where basically any challenges to copyright office decisions would need to be brought. For all intents and purposes, until the Supreme Court or Congress decide to weigh in, it’s pretty well established that AI generated content cannot receive copyright. There will probably be some decisions around the edges regarding when human editorial input might result in a copyright being awarded to the editor, but it would be very surprising to see the precedent overturned completely, particularly given its grounding in previous caselaw.
Authorship is also required in the Constitution's section on copyright, and is likely entangled in international treaties. Any change that foundational would probably require a constitutional amendment.
Did anyone expect it to go the other way? Europe is more hardcore about copyright than the US, and that's saying something. Germany is practically...well I was going to say they're nazis about it but perhaps that's in poor taste.
It's not surprising, though, since the ECJ held in the Infopaq case that copyright (under the EU Directives) can only exist for something which is the intellectual creation of an author.
> When you put your name on an article that's written by ChatGPT or Claude, you're basically putting a provenance mark on it saying: I take responsibility for this. I haven't written it, but I'm putting my name on it. That doesn't give you copyright, but it does give you liability for the content
If the AI is treated as an agent during inference, distinct from its user and not merely as a tool, then this should also apply during training too.
Based on this, it seems that AI agents are consuming people's code without permission. The MIT license only gives rights to "any person obtaining a copy of this software".
So the rights are given to a 'person', and the rights pertain specifically to a person who performed the act of 'obtaining a copy of this software'.
MIT license says 'obtaining a copy' and uses the word 'software', not 'code'. 'Software' is to 'code' what 'shop' is to 'building'; if you bought the shop, it doesn't necessarily mean you own the building. These are two different things and require different clauses. MIT explicitly separates the two and emphasizes that the author of the software retains copyrights (presumably over the code as this is the only thing over which they could claim copyright).
The code is different from the software; you can write the exact same software which behaves in the exact same way using completely different code; can be poorly written or well written. The difference is extremely meaningful to the person who invested effort to write the code in a clean way.
If we say that the agent is a separate entity from the person who ran it during inference, surely the same distinction can be made concerning the person who ran the agent during training. So the term 'person' from the MIT clause doesn't seem to apply here since the person running the agent has been factored out (just as they were during inference). Also, the agent is not obtaining a copy of the software; it's obtaining copies of the code which is copyright and independent of the software (as the MIT license clearly asserts).
Thank god! A victory for everyone who believes in the freedom of information, regardless of what you think about AI.
Copyright didn't always exist, nor should it continue to. Hell; it must not.
I think the words (read: hilarious 1.25pp pamphlet) of Aaron Swartz on the topic are just too poignant to ignore, given the paths of Reddit (corrupted yet democratic), IP law (malignant yet showing cracks), and government survellience have taken in the Trump era. Despite the dated context... he really says it best:
> I mean the copyright has to belong to somebody right?
Why would it? It’s generated by blending together ~ every bit of content on the internet and in books that they could steal. Why would the operator of the blending machine suddenly get copyright?
Does a gambler own the copyright on the symbols generated by a slot machine?
There is the potential for a solid piece of artwork here, ie. at what point does a string of words become an artistic product.
Imagine a slot machine in a gallery that produces a 4 word sentence. Visitors press the button, and the artist copies the 4 words onto a piece of paper and puts it on the wall.
EDIT: my point being that causing a machine to generate words doesn’t give you copyright over the words but if you then do something with the words they become your work.
If copyright does not protect AI-generated content, then AI should not be restricted from generating content that falls under copyright protection; yet, the reality is otherwise.
That’s not how that would work at all. The copyright of a derivative work has no effect on the existing copyright of an underlying work.
If you record yourself reading a book, you own the audio recording copyright but it would be a copyright violation to reproduce that copy without a license for the underlying rights.
In this situation:
If the AI generates the audio recording of a book, no one owns the copyright of the audio recording but it would still be a copyright violation to reproduce that copy without the underlying rights.
I'm not sure that logic holds. If I create a frame for frame copy of Star Wars then I cannot copyright it. However, that doesn't mean that I can still distribute my copy.
> If copyright does not protect AI-generated content, then AI should not be restricted from generating content that falls under copyright protection
What in the world makes you think that?
What is this implication based upon? Where does it say in copyright law that if no copyright arises for your work then it does not infringe the copyright holders' rights? This is completely devoid of logic...
This matches historical precedent in other nations where copyright was not extended to a photographer when a primate took a photo of themselves.
https://en.wikipedia.org/wiki/Monkey_selfie_copyright_disput...
All this talk of animal rights, and then this! /s
Yeah not funny.
Haaaahahahahah this is the smart humor we all come for here lmao!!!!
That was a terrible ruling. It should be about who got the camera there and set it up, not who pushed the button. With automated recording (dashcams) there isn't even "a button."
Hopefully no judge uses that as precedence.
>It should be about who got the camera there and set it up, not who pushed the button. With automated recording (dashcams), there isn't even "a button".
I'm actually having a bit of trouble thinking of what sufficient societal good there is/would be in granting copyright on raw dashcam or security camera or the like footage? None of those purely mechanical automated systems need a subsidy or encouragement to generate more. Certainly someone can use that sort of thing in the creation of a copyrighted work but what would be the issue with the underlying material in that case being unprotected?
I am in favor of copyright reform so eventually it becomes something like two years automatic with a one time two year extension possible an I agree with you because of one key word — raw. Once the footage is used in a package, be it a movie, a news report, or a music video, that package can be copyrighted. The problem is there is no way the owner of a dashcam can prohibit commercial use of a video they own if it cannot be copyrighted in the US. In a way, we have conflated copyrights with other so called intellectual property (details). Maybe the solution is to strengthen privacy rights somehow? I'm not sure.
Details https://ftp5.gwdg.de/pub/gnu/www/philosophy/words-to-avoid.h...
people freqently spend more than two years trying to get something published.
Not on YouTube.
Thinking more about monkeys, I think that case was made relatively easy because there was a normal living recipient for copyright, but that wasn't eligible for such. I'm fairly certain that dashcam footage, which is 100% automated, as you mention, is copyrighted by default. It's a fairly common source of footage of novel events like a meteorite, rocket impact, or whatever else. And that footage is licensed to media companies.
In any case, any country that decides neural network generated media isn't copyrighted then faces a major, and probably impossible, problem of trying to prove it wasn't made by a human.
Theoretically, in the US dashcam, CCTV, and similar automated recordings are not copyrightable. In the US you can only copyright creative expression. However, the question hasn't been answered in court, yet, AFAIU.
That doesn't stop people from treating automated recordings as if they were copyrighted, though. Especially for media companies, licensing is standard operating procedure. Even if a media company's lawyers believe something isn't actually copyrightable, if the licensor and licensee believe it is, and especially if distribution outlets (e.g. YouTube) treat it as such, then it all works out.
In some other countries automated footage is copyrightable, AFAIU, but it varies.
> In any case, any country that decides neural network generated media isn't copyrighted then faces a major, and probably impossible, problem of trying to prove it wasn't made by a human.
In the US the initial burden is on the person claiming a copyright violation. In court that is met by simply showing a certificate from the Copyright Office. In fact, it's actually a requirement--you can't sue in court to enforce a copyright claim with a certificate. But in getting the certificate you would be making a declaration that it was made by a human.
Copyright is automatic upon creation.
Registering a work provides additional benefits.
A dashcam could be a safety tool for most people, but it could also be a strategically placed B-cam for a videographer on a shoot.
For what it’s worth, dashcam footage can absolutely be copyrightable.
The ruling is more about “only humans can get copyright protection”, not so much anything about whether a button is pressed or not.
There is a spark of creativity in the monkey case that very often doesn’t apply to dashcam footage
Somebody exerted effort.
That’s it.
What if it was a human who took the photo, would the copyright still be owned by the equipment owner / person who set it up?
What if this applied in a photography class? The instructor owns the equipment and helped “set up” the photo. Does the instructor own the copyright?
After this ruling, I had an idea for a photo where I set up the entire set, camera, etc, but the photo entailed the model clicking the shutter while we were both in the shot. Just in case, I had the model write a quick note ceding the rights.
It's fine, I guess. How does it work in cinema? A director who is the creator of the project must have to get rights from every camera/mic operator.
> A director who is the creator of the project must have to get rights from every camera/mic operator.
No, because they're already paid to do that job. And a director is paid too - he doesn't have any copyright for his work. The employer does.
I think the issue you then run into is imagine another human came, picked up the camera, and used it to take a picture of whatever. It'd be fairly self evident that that the photographer would own the copyright to that work, but in your take - it'd be the camera owner.
So you're more of trying to create a special rule where if the normal recipient of a copyright would be invalid, then it slides to the 'nearest' most appropriate individual, but that seems extremely fragile and difficult to define.
Dashcams and security footage definitely don’t have copyright protection. Are you seriously thinking that the construction worker who set up the camera would then have copyright over everything it records?
Really? Regardless of the debate on what should and shouldn't be copyrighted, in the current system there is no way to set up a continuous camera with intention to create something and not be afforded the same rights as if I took a snapshot?
As software, creative works, science, etc become more and more contributed via AI does that mean all future works will not be copyrighted or patented? Under our current copyright culture and laws obviously not but that does open the question to how much does a human have to contribute and what evidence is required to show that a human contributed enough so that it can be copyrighted and patented. Some time in the future (may be many years) science may become to complicated for humans to understand fully with AI systems researching themselves. Current patent laws in many places including US says inventions created entirely by artificial intelligence cannot be patented. This future may be coming. What will the new copyright and patent laws look like in the future? Do we need copyrights and patents anymore?
It means as long as you can find a few employees who testify that they made "creative input" during the process and nobody can disprove that, you're fine.
It's like things that are already in public domain. Even if you make a coloring book out of paintings in public domain, it doesn't necessarily mean others can just print your book as-is.
> It means as long as you can find a few employees who testify that they made "creative input" during the process and nobody can disprove that, you're fine.
Have you been involved in copyright or patent litigation?
Because it's not that easy.
It's usually easier than that, as one by default hold the copyright of their works. Typically a company doesn't need to provide the whole working process and files to the court to claim copyright over their published products / to accuse pirates for violation.
If you're on the defendant side of a copyright violation case, it's extremely hard to use "well the original author didn't really make it...* as a defense. (Patent cases are often defended with this argument though, as a patent grants far boarder protection than copyright and can be rejected on prior art. But still it's very different from "AI made this actually.")
> It's usually easier than that, as one by default hold the copyright of their works.
That's not what we're talking about.
This is a question of copyright eligibility in the first place. An excellent recent(ish) example is Oracle v Google wherein the former claimed copyright over the Java APIs. That particular litigation went on for years and ultimately went to the Supreme Court (though it ultimately ended on a fair use argument and the issue of API copyright eligibility was never settled).
We can probably assume similar high profile cases in the future that turn on quantifying the level of human contribution in code produced through the use of AI.
As for patents, it's literally that easy. Attacking authorship of a patent is an excellent way to invalidate it. It's why any company with a serious patent program requires staff to keep supporting notes for their work and confirm authorship in writing as part of the filing process.
You can blame the governments for all sort of inefficiencies but any operation that impacts taxes runs a tight ship in most developed nations, this include things like copyright law, census, and anything statistics.
If companies get slope with creative output to the point that "a few employees" can reproduce it in shadow markets, don't expect to get copyright protection without giving governments revenue and speech-control.
Perhaps but what about individual works done not via a company. Say you are a author and someone falsely detects a AI watermark in your book. Is the testimony you provide enough to override that falsely detected AI watermark? My feeling is that there will be an industry that arises that proves you contributed to the works you want copyrighted.
(I think) the case depends on what AI watermarks are. As far as I know, things like SynthID has practically zero false positives.
My understanding is that you get back a percentage from SynthID regarding how confident it thinks the text is generated by AI or not. The question is what percentage would be considered undeniable that it is AI generated or require more than just your testimony to prove that you actually did the work yourself. https://www.nature.com/articles/s41586-024-08025-4 (Google Deepmind) was using a "Watermark detectability is measured using the true-positive rate (TPR) when the false-positive rate (FPR) is set to 1%" for instance in their research paper. This 1% was set arbitrary just for the research paper but what percentage should it be set to?
I think you need to provide independent sources for the "SynthID has practically zero false positives". Except a paper written by Google authors about SynthID for images I have not found any study about zero false positives.
Like being an executive producer
"Content that is entirely generated by artificial intelligence is not protected by copyright."
If that is faithful reading of the law, that makes sense. I know a number of people who use AI, but none of them (that are making anything actually useful) have the output "entirely generated" (aside from some POC tests that never see the light of day).
I have a hard time believing anything of value, anything worth copyrighting, could be entirely generated by AI.
The current trend (since months now) is to have everything done through agentic loops. Meaning humans are only here to provide the initial prompt and do a few back and forth during the implementation.
> I have a hard time believing anything of value, anything worth copyrighting, could be entirely generated by AI.
Perhaps you're not stretching your imagination enough. What if an expert novelist used an AI like it were a fancy auto-completing dictation machine to write the next great American novel? AI may have "entirely generated" all the text, but what if they micromanaged the shit out it?
I can imagine the difference between someone who fires off a lazy 5 minute prompt, and someone who labors for months and months to get exactly the results they want.
It’s almost like laboring for months and months would have draft revisions, dozens of changes, and many other interactions they could prove as theirs defeating your entire argument.
This is less relevant for logos & creative works, but things that enable companies to do production will go back to being closely guarded and sharded secrets, which is what the patent system was trying to resolve (though we can certainly argue the success of it).
If companies fail to protect their investments in generating IP, they will stop investing in generating it.
And unless IP generation costs (all in, including the humans telling them what to generate) fall close to zero, it will be bad for the world if companies cannot recoup investments in generating new IP.
We would expect this to hit those industries relying on IP protections the most, e.g. pharma.
> Content that is entirely generated [...]
Take note of the qualifier entirely. If you're working with an agent steering it to produce the results you want, it would be an entirely different story.
Wheres the limit. Would a prompt "Create a system that does x" and no other feedback count? Or do need to do 10 prompts to steer your agent to the results you want. What about just 1 additional prompt. Somewhere there will be a line.
Legal systems are fuzzy. The limit will be what courtrooms decide. There will never be a ">75% of content" law, everything will be left to interpretation, laws are not source code.
a bit further down ...
> Neither mere prompting nor the selection between several AI suggestions is sufficient as a human creative contribution.
which reads like you actually have to contribute by modifying the output from the system, i.e. prompts/inputs to the system do not count as a human contribution.
If you modify the output, you don’t magically get copyright on the whole output, only the section you modified
Is working with an agent different than constantly editing an image with fine tuning prompts? Both can be entirely generated.
Where is this qualification in the legal language?
Where is the legal language?
Exactly, I don't know how the person I was replying to was making such confident statements about the situation.
Why would copyright and patents exist when the US blatantly stole the entire world's knowledge to train their AI models (and China distillled from it)
Knowledge cannot be "stolen"
I think that you might be taking "knowledge" too literally. As I read it, the parent's post was using "knowledge" poetically for the various works that AI companies stole.
What about art?
You wrote your house's security code on a Post-It note. Somebody came in as a houseparty guest and located the note, and discretely copied it onto their hand to get access to the house during your next vacation.
Which verb would describe their knowledge of this security code?
Obtained, read, viewed, learned, discovered.
The knowledge wasn't obtained illegally or illicitly, immorally perhaps.
Illegal entry/trespass is a crime regardless of the method used however.
In your analogy copying the code (online content) isn't necessarily illegal, however using that knowledge without compensation, permission or attribution may be a crime (depending on copyright law, judges rulings etc).
https://blog.curiousquail.com/im-upset-again-about-a-co-crea...
What about using torrents to acquire that content?
Torrents are legal?
Seeding is often considered the actual point of infringement, aka sharing copyrighted material without the right licence.
I should note I suppose that I don't agree with companies using artists or developers hard work for their own financial gain without compensation, whatever form that may take.
The problem as usual is the availability of legal loopholes when you have a team of lawyers and near infinite money.
It really depends the country, downloading a Hollywood movie via torrent is itself illegal in a lot of countries. The only exception I’m aware of is Switzerland
You're right I was thinking of civil vs criminal penalties.
Personal downloading is generally only a civil violation, requiring the rights holder to sue you for reparations.
This point makes sense to me. I agree with "learned" and "discovered".
"eavesdropped"?
"spied"?
Whatever it is, slapping "steal" on it because what was done is wrong, and stealing is wrong, is the same sloppy logic that the copyright lobby used in the 80s/90s.
If 'steal' doesn't fit, then neither do 'eavesdrop' or 'spied'.
The crime there is using that knowledge illegally to break in.
learned
After a bit of thinking, I agree that "learned" is the most appropriate. I stand corrected.
So this means you can not use AI for the majority of open source projects since GPL, MIT, BSD, etc are all copyright declarations and they'd be being made for code which you can not copyright.
You can't enforce those licenses against AI-generated parts, because enforcement relies on the recipient having no other way to avoid copyright infringement.
However, there's no issue with including non-copyrightable code in otherwise copyrighted projects. There's already plenty of non-copyrightable code like auto-generated boilerplate.
>no issue with including non-copyrightable code in otherwise copyrighted projects
That is not how copyright/trademark/contract laws work, and isomorphic plagiarism is not a long-term business model. People also loved Napster at first too. Good luck =3
https://www.youtube.com/watch?v=YhgYMH6n004
No it would mean the license does not apply to the parts of code written by AI.
I find it hard to find correct terminology in this case. AI generated content is copy right wise in state of nothingness. It simply does not have copy right status like other material can have. As such maybe best I can formulate is that you can't enforce license violations against ai generated material as you do not have ownership.
So no license is enforceable with code written by AI.
No, GPL is a contaminating license contract, so "AI" slop means you are probably in GPL violation by including isomorphic plagiarized uncopyrightable code. =3
wasnt there a case a while back, where GPL or LGPL code made its way into MIT licensed software via LLM? And they were forced to remove the copyleft code. I dont remember the details though.
Usually what I have seen is someone writes a Perl library, Claude mostly ports it to Python with isomorphic plagiarism, and a bunch of Brogrammers proclaim "AI" magic is real.
The issue is most GPL license fall under contract law, and scraped code can't legally have assigned "copy" rights on an "AI" vector search compaction output.
https://www.youtube.com/watch?v=YhgYMH6n004
Indeed, these rules obviously don't apply in places like India, Russia, Iran, and China. =3
Indeed, but people will continue to think vector search compaction similarity absolves folks of Trademark and Copyright liability.
As the dark specter of Disney Mickey Mouse looms over every LLM model involved in isomorphic and character plagiarism. Yes, even motion capture is considered a performance act in the guilds, so video reskinning an unlicensed performance act people make is also a liability.
It would sure save a lot of money if you don't get caught, so people are gonna try it for sure. =3
It will be interesting if or when someone tries to enforce that perspective. Good luck proving an online contributor used AI on a merge request.
Perhaps we'll have new iterations of FOSS licenses to adjust to legal declarations.
We know for a fact that Bun port to Rust was done entirely by Claude! I’m wondering if that could be used against Anthropic in some ways (not that I want to, just curious what would be the angle)
We’re soon going to return to this: https://en.wikipedia.org/wiki/Trap_street
Add a subtle bug or odd behavior to your code, see if someone else’s code repros it (presumably via LLM regurgitation), sue.
It is actually pretty easy when the area is very specific like 1 guys Perl library, and Claude ports it to Python for a group unaware of what happened.
All models know what Disney Mickey Mouse looks like too. =3
This post is glossing over most of the nuance in EU law.
The AI system must function merely as a tool or instrument (like a camera or Photoshop) guided by the human, rather than acting as the creator itself. The line may get a bit fuzzy case-by-case, but effectively the human must be the creative one, not the AI.
This is not unprecedented. Machine generated technical data, sensor outputs, automated surveilance photography, monkey selfies, purely algorithmic or generative music and such were already disqualified long before AI came along.
Interesting thought experiment is to consider an author who writes a novel in English, but wants to translate it into German.
They have a copyright on the original, and if they hire a human, the human would have a copyright on the translation (which would generally be licensed or transferred back to the author in some way).
If they use an AI for the translation, by the logic here, the translation wouldn't have its own independent copyright, but (based on other long established principles of copyright) it would still be a derived work of the original, so even if this decision holds it would not be legal to make unauthorised AI translations, pirate authorised AI translations, make further translations into other languages (or back to English), etc.
Which seems fairly reasonable! But consider:
If you start with, say, a 90,000 word novel, and ask for a translated novel, you (presumably) have sufficient rights to stop someone making unauthorised copies of the AI translated version.
If you start with a 300 word prompt, and ask for a logo, you (apparently) do not have sufficient rights to stop someone from using it without authorisation.
So some combination of the input (0.3k vs 90k) and the output (logo versus novel) crosses an inflection point between these two extremes, and I think it's interesting to wonder what the boundaries are. Like, in theory you could graph input size versus output complexity, and sketch a frontier between "the author's protected expression survives in the output" and "the author's protected expression does not survive in the output". And I don't have the slightest idea what I think a fair frontier would look like.
> Some combination of the input (0.3k vs 90k) and the output (logo versus novel) crosses an inflection point, and it's interesting to wonder what the boundaries are
It has always been like that though. Copyright is just that arbitrary. You really can't tell if Android violates Oracle's copyright over Java by reading law text.
It's supposed to be arbitrary. Copyright is not some sort of human right or divine commandment that creators have. It's a legal device to incentivize a certain type of behavior. Like tax cuts and subsidies, governments can dangle it in front of people whenever it wants.
Monkeys and LLMs (so far) either don't understand or don't need such incentives, so they don't get copyright.
All law, in fact, is arbitrary.
I've come to understand that rules are not executable, or not executed. As programmers we understand if we write if(X) {Y} and X happens, Y will.
Laws aren't like that. Laws are just powerful people telling us what we want to do and what they want to do to us if we don't. The difference is subtle: the powerful person's power is not limited by the law. He only writes the law because exercising his power is annoying and he'd prefer that we complied in advance so he doesn't have to. The thing that actually executes is the powerful person's desires, not the law.
Humans dont need incentives either, we an abolish it.
Individual humans, maybe.
The Walt Disney Company insists otherwise.
Fuck Walt Disney
Input size and output complexity are not necessarily the only two dimensions involved. There’s no reason to think that such graph would show a continuous frontier.
I don't see why the author of the 300-word prompt couldn't claim copyright of it.
As far as I'm following it, the court didn't suggest the prompt was not a protected work.
But what the court did say is that the logo wasn't (sufficiently) derived from the prompt; those 300 words weren't contained or encoded in the logo, so if you copied the logo without permission from the author of the prompt, this was fine, because you weren't in any meaningful sense copying that prompt, or any other protected work of the author. (Which in my view is fair enough, really!)
Another way of thinking about it is to consider:
1) The number of different prompts that would yield a very similar looking logo. In many cases, a one or two word difference would probably yield a very different logo. (Conversely, you could rewrite an entire chapter of the novel, but the other chapters would still be translated more-or-less unchanged. You'd need to rewrite every sentence to end up with an entirely new set of sentences in the output, and since copyright covers the arrangement of text, not the plot or ideas, that matters.)
2) How feasible it'd be to look at the logo and accurately guess the prompt. (Again considering the hypothetical novel, if you translated the novel back into English, entire passages would likely be word-for-word identical with the original.)
it's sort of non-controversial at the extreme end I think - where there is no human contribution, why should there be protection?
But it just begs the actual question of how much human contribution there needs to be:
- I wrote the prompt (not enough)
- I wrote many prompts and iteratively refined them using distinctly human skill (open question, but loosely seems still not enough, potentially in the EU but maybe in the US?)
- I made minor modifications post-generation (open question, probably enough)
- I made equal or more contribution to the final result (this better clearly have copyright protection or we are in real trouble)
> I made minor modifications post-generation (open question, probably enough)
Hmm I don't think that would be enough. I'd expect that would only make the modifications themselves copyrightable, but not the whole modified work including the AI parts.
Compare for example the case where the US copyright office ruled that, when assembling AI-generated images and human-written text into a comic book, only the human-made elements themselves (text, arrangement) get copyright protection, but not the images.
E.g. if you generate an AI photo and color grade it, i expect only the color grading would be protected (if that is even significant enough to be protectable), not the rest. And someone else could re-color-grade the same image without infringing your copyright.
I'd say that any use of an AI trained on people's creations without (real) consent and compensation should not be protected.
Is it fair to say, then, that artists who have trained on the work of prior artists should also not be able to protect their works unless they have directly received consent and have compensated all the artists whose works they have viewed and potentially learned from over their lifetime?
> AI generated logos do not enjoy copyright protection
But they can still be registered as trademarks.
Does this also extend to generated source code? I mean, this upend the business models of many A.I. companies like OpenAI and Anthropic. In addition, many ISVs would probably stop using A.I. if their products could lose their copyright.
And what about A.I. generated code which has been altered by a human? Is the derived work copyrightable?
And this is one of the reasons some people are up in arms about Anthropic's recent inclusion of watermarking against anything Claude models touch. It will even take human-generated code that it works on and apply watermarks to it which can be then used to prove it was AI processed and perhaps invalidate any copyright on even human-written code.
Does this mean someone could develop a switch emulator using an AI and Nintendo couldn't win a lawsuit against them?
Does it enable decompilation remasters of classic games?
It feels like AI is a cleanroom laundromat
Im not sure it’s really a copyright issue here, and not patent or trademark related, or something else. Copyright is really only one aspect of IP laws and they all have their own legal nuances
Not a lawyer, but I'd expect that decompilation creates a derived work which is still subject to the copyright of the original.
If you engineer a clean room process, it does not. Already proven with an open source project that went from GPL to MIT via clean room (chardet).
No, an AI's output can still violate copyright.
Definitely not clean room, quite the opposite.
“That doesn't give you copyright, but it does give you liability for the content.”
I don't see an easy test here. Worse, I see the beginnings of a test that is technical and very likely to not match the future of how we will interact with these things. We have to start somewhere but I'm not sure 'Neither mere prompting nor the selection between several AI suggestions is sufficient as a human creative contribution.' is the right place to start. I think we need more examples of what does constitute 'human-centric' and work from there. I also don't think that any system that says 'you didn't do enough work so it isn't human-centric' works. Any system like that will require a reexamining of copyright in general. There is a lot of low work copyrighted material out there. Just because 'AI' didn't build it doesn't mean the same tests shouldn't apply.
Some examples of why I think this is really hard: Say I build a story generation system. I work hard on building an agent swarm of actors, critics, editors, researchers. I craft into the various agents concepts of story arcs, outlining techniques, character development. I build a huge well thought out process for how to agentically write an actually good story, so long as you give it a title. Heck, I even design and train my own custom LLM with original layer ideas and novel training techniques to use on this system. After all that I then take that final step and give it a title. Do I have no claim to that? I probably put more work and creativity into it than an author would have a book. What if I then gave it 500 titles? 5,000? Would my claim degrade the more titles I fed it? Is it a percentage of work question? What is the core concept here that defines 'human-centric'? What is the cut-off here?
Let's go even further. I don't prompt. I live in a world with unlimited context models. I have a conversation about the book I want it to write. During that process I reject some ideas and accept others. I didn't give it a 'system prompt' but essentially all I did was prompt it and select versions I liked. Is that not human centric? How about if I asked it for advice and it did some editing work on my story? Did that make it not human centric even though the starting text was mine? What if that starting text was 99% replaced with a version 10x as verbose. Defining based on how you interacted with the model (prompted and selected) just seems way to weak to be a clear test.
I think there is an easy test: if you didn’t write it, it’s not yours.
If we go back 10 years and your friend says “I have an idea for an app, here it is,” and you build it, you own the copyright because you wrote it.
You give an idea to the pile of math calculated of the stolen work of humanity, the math owns it (which it can’t, so no one owns it).
No matter how detailed of a conversation you have with a friend, I don’t think they have justification to claim copyright over code written by you.
That's how you want it to be, but that's not how it is.
Actually, in Germany, copyright remains with the author and their heirs until seventy years after the author's death; it cannot be sold or given away, even if someone paid them doing it. Instead, there is "Verwertungsrecht" or "Nutzungsrecht" ("License", "Usage right"). In the USA, copyright can be sold completely to another legal entity.
Then how it is needs to change.
What if an artist programs a programs a carving robot to create a sculpture? How is that meaningfully different?
Based on your logic that should not qualify, but it currently clearly does: https://quayola.com/selected-unfinished-sculptures/
I think the copyright should go to the tool creator in that case, if the tool is acting as an extension of the creator.
But if the tool is created from collective human creation, the copyright should belong to all humanity, not the person who triggered the tool.
If you trained an LLM entirely on your own input, I think you should own the output, but that is not the case for any widely-used llm.
Information wants to be free. We should simply dispose of copyright. With LLMs, waters that were already muddy are now a thick slurry. And it's only going to get worse.
It's an antiquated mechanism and is far more abused than it is actually used at this point.
If we got rid of copyright, what would prevent corporations from simply stealing artist's work?
Isn't that what the AI companies already did?
My test would be: Until the AI companies pay back what they have already stolen, then anything that uses AI at all cannot be copyrighted.
Who set the value? I wouldn’t have sold any of my content to ai companies for training, the payback would be basically infinite
Lots of wild guesses about mixed human-AI copyright in here. Last time I read the copyright office’s statements in this, their stance was basically: The human owns exactly what they contribute. The rest is public domain.
Yes, that is vague. I think the examples were like:
If you paint a symbol and use an AI filter over that to stylize it, you own the symbol aspect of the image but not the stylized final result.
You can own a book of AI images as a curated collection. But, not the individual images.
Seems about right. Copyright is supposed to literally prevent outright copying. The output from an LLM is not a creative work of the prompter. It's genuinely the opposite, I use them when I don't care about something but have to do it anyway for whatever reason. It makes more time for me to do the things I like working on.
It feels shitty that they have been trained on the life sums of all of our work and online presences with absolutely no credit given... But then again, I'm not sure I'd want to know what parts of the weights were from me and which weren't.
I feel like people will be upset by this take... But you don't get credit for the creativity of an idea alone. Every asshole has ideas. It's called a work of art for a reason, emphasis on WORK.
I’m trying to imaging what giving credit would look like… “Credit: Everyone who has made anything freely observable on the internet.” Would folks be less upset if that was included?
Frankly, I should have been putting something similar attached to any drawings or paintings I make physically. My entire life I’ve been training my brain on countless works of others. And, I’ve never asked one of them for permission or given anyone any credit.
1) Does applying a copyright message warrant that the content has not been generated by an AI?
2) If one applies a copyright message to AI generated output, is that fraudulent?
We've reached the stage where copyright has become impossible. Let's just accept it and move on.
Heck yeah.
https://fmhy.net
That's hell of a link sir.
If you want something specific fast I know a thing or too but it is all in the link.
There was some related discussions of US law recently[0].
[0]:https://news.ycombinator.com/item?id=49353603#49354813
Suppose that an artist trains an AI model purely on their own works, and then the AI generates something similar to that artist's work. I would say that the artist must be allowed to assert copyright over that. The artist has copyright over all the training data, and the output of the AI is a derived work of that.
What do you think should happen if an artist views a lot of arts over their lifetime and then themself generates something similar to another artist's work? I would argue that collectively humans have been standing on the backs of other humans all throughout history. An artist today does not simply come out of the womb knowing how to make great art - they learn it by observing other art, learning about art and practicing (not altogether dissimilar to model training) over many, many hours.
I really wish the « what about humans » line of reasoning would die, it’s a lazy way to avoid engaging with the actual LLM questions
Such situations have led to disputes which ended up in court.
Microsoft is using AI generated code in their OS, so does that mean they lose copyright protection in Europe?
Once it becomes too substantial and there is not enough human oversight in process.
Say I would prompt Unix like system in few prompts or started agent chain to make it. That OS wouldn't have a protection.
Does it matter?
The set of possible and desirable non-copyrighted text/audio/visual states to render is effectively infinite. Laws don't prevent Scrabble clones; tropes are not protected.
Endless remixes of public domain content are an option as well.
If an iPhone with model weights in chip, a local Mac mini with similar but more powerful model on chip ends up capable of generating endless content copyrights won't provide a moat.
Why should anyone want to have copyright on AI-generated content?
For the same reason they'd want copyright on anything else.
I feel like this question is typed wrong but the same reason someone wants copyright on human work to make money by prevent other people from stealing your work
Not just money, FLOSS relies on copyright laws
To be able to enforce your license?
To profit under intellectual property constructs. As it stands, AI output is effectively in the public domain.
Given the cost of generating images/text/etc. is much lower when using an AI system, as compared to a human, this seems appropriate. The purpose of intellectual property is to encourage people to invest that time into creating works that benefit others. Perhaps there’s an argument for a shorter copyright period to cover the cost of the compute spent, but certainly one wouldn’t expect AI intellectual property to require a similar level of incentive to produce.
Probably unless it infringes on others' copyright as it does...
Does a photo/image that generated/modified by a human-being using photoshop got protected by copyright? I think previously someone does this...
Year ago, there was a painter, Thomas Kinkade, the "Painter of light".
He became a bit of a pop phenomenon, with stores dedicated to selling high quality prints of his works.
And one of the things, so I'm told, that could happen, was that you could buy one of his prints, but someone at the store, an artist, would dab some paint onto the print. Add some "light" to it.
Obviously this is a commercial endeavor, so I doubt the provenance of the rights holder was in doubt (such as through an employment agreement).
But it's, perhaps, an interesting case study about who owns what in a time of augmented and manipulated media.
I believe those are called “embellishments”, and they’re semi-common as a way to make a print “unique”, but still cost way less than the original the print was copied from. From the couple artists I talked to about them, their students do these at their studios, rather than the artist themselves.
Yes it does because there's is still direct control from a human. When working with LLMs, even interactive prompting does not provide that level of control.
You can copyright the instructions file, as long as you wrote it.
This then implies that in the EU, AI-generated content cannot have a license applied? MIT, GPL, CC, etc?
This has already also been implied in the US. Courts upheld the Copyright Office's stance on human creation in the context of AI image generation. There's no reason to expect something that fundamental to copyright to be any different for other media, such as source code.
It's reasonable to expect this to hold for all Berne Convention countries.
All licenses are unenforceable if you don't hold the copyright, since you don't have a copyright to license. And attempting to do so would probably be perjury.
Yes. Also CC is complicated in EU. A creator can not "give up" their copy right. Work can not put to "Public Domain" as creator.
Is that for the whole EU? I understood it was country specific, for example France doesn’t let you give up your copyright, but other EU members do
Makes it even worse. Basically it hugely depends. For example Finland allows letting go of commercial exploitation rights and even make binding statement on that. But still there always remains right for them to be informed if copies are being published.
Not to even mention moral rights on how such work can or can not be changed if it would be disparaging for example...
IANAL, but yes, I'd assume the license is unenforceable on the AI content.
Under this interpretation models themselves do not have copyright protection either
Funnily I think copyright protection for LLM model is 15 years. As to me it just looks like very fancy database and that is the protection for those...
Does this apply to software?
yes, why wouldn’t it?
I will add a pixel by "hand" with software on my computer.
That’s not enough to generate copyright protection.
The creativity requirements may seem arbitrary but there’s a legal distinction between a sculpture and a standard brick.
>The creativity requirements may seem arbitrary but there’s a legal distinction between a sculpture and a standard brick.
Or more relevantly, a recipe find on recipe sites (with the author's entire backstory) vs a sequence of instructions. The latter is not copyrightable, even if there was some creativity that went into it (eg. word choice).
Judges are not at all stupid enough to fall for that. You would get laughed out of court for that. Judges are not computers, and they can not be tricked with these kinds of technicalities.
In fact, the defense of "I wrote the prompts that led to the code that the LLM wrote" would be a much better defense.
My bet would be to build anything from multiple independently generated pieces of AI code. Thus the software build from AI generated blocks by human would at least have protection on basis of the structure and work put into that.
You will have copyright of that single change
>Content that is entirely generated by artificial intelligence is not protected by copyright.
entirely is the plank supporting this.
Interesting. In addition to Anthropic's watermark use to prevent model collapse, we can definitively call something copyrighted or not copyrighted. This is a boon to everyone who consumes culture.
That doesn't actually work. Say you took a book and used an AI to translate it into another language. The translation wouldn't have an additional copyright, the way it would if a human had done the work, but the output would still be restricted by the original copyright. So the presence of the watermark does not tell you that the text is public domain.
I thought translations can be copyrighted separately. Their are translators for instance who translate very old texts from other countries/languages and then sell that book/translation using copyright to protect this business models? Not sure but does this mean if you as a human wrote a book in English and then used AI to translate to another language say French then that translated work would not be copyrighted. Not sure how this would work.
Translation, in and of itself, is viewed as a creative work. A translated work has two copyrights: the original, and the translation, with permission needed from all rightsholders to redistribute. A new translation of a public domain work (Emily Wilson’s The Odyssey) has one copyright holder, the translator. An existing work (the English translation of The Three Body Problem) has at least two: the original author and the translator.
However! Since AI work is noncopyrightable, the AI’s effort in translation is simply ignored. Claude’s The Odyssey would have zero rightsholders, and remain public domain. ChatGPT’s translation of The Three Body Problem would still be under Liu Cixin’s copyright.
Lmao. No. Content that is wholesale generated by AI is also not subject to copyright in the US. It relies on the honor system and you can always modify the material juuuust enough that you can claim the copyright. In practical terms, this changes nothing.
Good luck proving the origin of these created works
The same is true in the US. This is a nothing burger.
It hadn’t been established yet in the EU. It would be a big deal if it went the other way in Europe so it being the same isn’t a nothing burger.
It hasn’t been established in the US either, despite how desperately some people want to misinterpret the handful of cases so far.
It’s been established in the DC circuit, where basically any challenges to copyright office decisions would need to be brought. For all intents and purposes, until the Supreme Court or Congress decide to weigh in, it’s pretty well established that AI generated content cannot receive copyright. There will probably be some decisions around the edges regarding when human editorial input might result in a copyright being awarded to the editor, but it would be very surprising to see the precedent overturned completely, particularly given its grounding in previous caselaw.
Authorship is also required in the Constitution's section on copyright, and is likely entangled in international treaties. Any change that foundational would probably require a constitutional amendment.
Very clearly established.
AI only = no copyright.
“A Single Piece of American Cheese” got a copyright because it had human involvement in compositing.
Theatre D’Opera did not because it was primarily prompt driven.
Thaler didn’t because he said it was machine derived.
Humans must be involved for a copyright.
Did anyone expect it to go the other way? Europe is more hardcore about copyright than the US, and that's saying something. Germany is practically...well I was going to say they're nazis about it but perhaps that's in poor taste.
It may be a nothing burger for US companies/persons, but the EU is fortunately not governed by US law. This is big news for us here.
It's not surprising, though, since the ECJ held in the Infopaq case that copyright (under the EU Directives) can only exist for something which is the intellectual creation of an author.
It's not big news if you know anything about Europe in relation to IP law. No one expected a different result.
> When you put your name on an article that's written by ChatGPT or Claude, you're basically putting a provenance mark on it saying: I take responsibility for this. I haven't written it, but I'm putting my name on it. That doesn't give you copyright, but it does give you liability for the content
If the AI is treated as an agent during inference, distinct from its user and not merely as a tool, then this should also apply during training too.
Based on this, it seems that AI agents are consuming people's code without permission. The MIT license only gives rights to "any person obtaining a copy of this software".
So the rights are given to a 'person', and the rights pertain specifically to a person who performed the act of 'obtaining a copy of this software'.
MIT license says 'obtaining a copy' and uses the word 'software', not 'code'. 'Software' is to 'code' what 'shop' is to 'building'; if you bought the shop, it doesn't necessarily mean you own the building. These are two different things and require different clauses. MIT explicitly separates the two and emphasizes that the author of the software retains copyrights (presumably over the code as this is the only thing over which they could claim copyright).
The code is different from the software; you can write the exact same software which behaves in the exact same way using completely different code; can be poorly written or well written. The difference is extremely meaningful to the person who invested effort to write the code in a clean way.
If we say that the agent is a separate entity from the person who ran it during inference, surely the same distinction can be made concerning the person who ran the agent during training. So the term 'person' from the MIT clause doesn't seem to apply here since the person running the agent has been factored out (just as they were during inference). Also, the agent is not obtaining a copy of the software; it's obtaining copies of the code which is copyright and independent of the software (as the MIT license clearly asserts).
Thank god! A victory for everyone who believes in the freedom of information, regardless of what you think about AI.
Copyright didn't always exist, nor should it continue to. Hell; it must not.
I think the words (read: hilarious 1.25pp pamphlet) of Aaron Swartz on the topic are just too poignant to ignore, given the paths of Reddit (corrupted yet democratic), IP law (malignant yet showing cracks), and government survellience have taken in the Trump era. Despite the dated context... he really says it best:
https://ia800101.us.archive.org/1/items/GuerillaOpenAccessMa...
Create an “assign copyright” skill?
I mean the copyright has to belong to somebody right?
> I mean the copyright has to belong to somebody right?
Why would it? It’s generated by blending together ~ every bit of content on the internet and in books that they could steal. Why would the operator of the blending machine suddenly get copyright?
Does a gambler own the copyright on the symbols generated by a slot machine?
There is the potential for a solid piece of artwork here, ie. at what point does a string of words become an artistic product.
Imagine a slot machine in a gallery that produces a 4 word sentence. Visitors press the button, and the artist copies the 4 words onto a piece of paper and puts it on the wall.
EDIT: my point being that causing a machine to generate words doesn’t give you copyright over the words but if you then do something with the words they become your work.
The argument could probably be made that we currently don't recognize AI as a person and therefore no one created the work.
There does not have to be copyright. Nothing mandates that piece of data has one.
If copyright does not protect AI-generated content, then AI should not be restricted from generating content that falls under copyright protection; yet, the reality is otherwise.
That’s not how that would work at all. The copyright of a derivative work has no effect on the existing copyright of an underlying work.
If you record yourself reading a book, you own the audio recording copyright but it would be a copyright violation to reproduce that copy without a license for the underlying rights.
In this situation:
If the AI generates the audio recording of a book, no one owns the copyright of the audio recording but it would still be a copyright violation to reproduce that copy without the underlying rights.
These two things that are quite distinct. Whether something is protected by copyright, it should not be counterfeiting regardless.
I'm not sure that logic holds. If I create a frame for frame copy of Star Wars then I cannot copyright it. However, that doesn't mean that I can still distribute my copy.
> If copyright does not protect AI-generated content, then AI should not be restricted from generating content that falls under copyright protection
What in the world makes you think that?
What is this implication based upon? Where does it say in copyright law that if no copyright arises for your work then it does not infringe the copyright holders' rights? This is completely devoid of logic...