Rendered at 07:16:57 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
SillyUsername 1 days ago [-]
If they're going to treat copyright works like a public good, then the product they produce should also be a public good to prevent the "free-rider" problem.
---
The free rider problem is also a form of market failure...
The production of public goods results in positive externalities which are not remunerated. If private organizations do not reap all the benefits of a public good which they have produced, their incentives to produce it voluntarily might be insufficient.
---
The US government has just made the NY Times content a public good, and in doing so, allowed openai to create a private good, for which they will be remunerated instead.
It could be argued that this means the US government has effectively taken profit from one company to fund another, under the guise of helping global competitiveness, and whilst this may be the current M.O. of the Trump administration it means predominantly domestic companies are effectively forced into providing export subsidies for the new industry, and monopolising (or oligopolising?) an incumbent.
I wonder if on this basis Seedance could also be equally legally be allowed to use copyrighted media for its video AI, so effectively they're also pulling the legs out from under the film industry too.
Venn1 1 days ago [-]
I won't be the first on HN to mention how having a niche tech blog is becoming unmaintainable, and this reads like a greenlight to scrape away. The Creepy Crawlies post[1] from a few days back was basically a checklist of nonsense I have to deal with, albeit on a smaller scale. Even so, it's getting outside my budget.
I'd much prefer a world where everything is not locked behind a login with bonus 2FA, but that sure seems like where we're headed.
This article is trash, but this ruling is just about the outputs of AI.
For use as training data you still need some right to it AFAICT.
m4rtink 23 hours ago [-]
So copyright is now effectively dead ?
general1465 20 hours ago [-]
As long as you transform copyrighted data into something else then it appears so.
Which will be also interesting for prompts used by the users of these AI companies, because even if those prompts contain copyrighted data, they can be used for training because training is transforming these prompts into something else.
asksomeoneelse 20 hours ago [-]
Only for the work of the plebeian. They are still very much in effect when it benefits our overlords.
conartist6 1 days ago [-]
What's mine is yours, comrade!!!!
trimethylpurine 1 days ago [-]
The CCP adopted capitalism too.
There's no sense in worrying how closely the university aligns real policy with various theoretical economic ideals.
If we allow copyright law to directly inhibit technological advancement then we defeat the very purpose of copyright law.
That said, it need not inhibit if they just pay for using it. Hopefully they will sue for reasonable compensation, noting that they have such a privilege only in a handful of countries so any restriction is strictly a restriction on those countries, which is a huge technological advantage for all the rest.
I try to see both sides.
doc_ick 11 hours ago [-]
Copyright law can be misused and can inhibit some advancement, but the fact that “certain groups” or “certain companies” can just ignore laws with no ramifications is something else. If it’s probably life-saving or a pipeline of a dream (see Elon musk always promising FSD), then they should pay or prove the end result.
trimethylpurine 43 minutes ago [-]
I agree. You can be right (in my opinion), and stupidly file a bad case (in my opinion), simultaneously.
You can be right, and dead at the same time.
Just because I think they filed stupid doesn't imply that I think they are wrong.
conartist6 23 hours ago [-]
Yes, just give me my food ration comrade. It is all I need to keep making art for the glory of The Party and it's Supreme Leader. Everything I do I do for the glory of AI. AI is progress!
trimethylpurine 41 minutes ago [-]
Personally I worry that AI is an existential threat that is totally out of anyone's control by now.
Not sure if you read my comment in a different light or not, but I do get the feeling you misunderstood my meaning.
DarmokTanagra 1 days ago [-]
[dead]
BottieZimmie 1 days ago [-]
[flagged]
kittikitti 1 days ago [-]
The New York Times prides itself as staying relevant with digital content. Forced trends like Wordle are still staples in their marketing. It's ironic that they continue a smear campaign against artificial intelligence which includes suing AI companies for training on their digital content.
I would like clarity on copyright rules and regulations but it's hard to sympathize with NYT when they have been so aggressive with their AI fear-mongering. The authors and artists must be the centerpiece instead of giant corporations. I doubt they will use any leverage to actually improve ethical usage of copyright with AI, and they only want to protect shareholder profits.
I think it makes sense to say training on material you legally acquired is fair use. Copyright, quite famously, doesn't protect ideas. Nobody really needs or wants AI models to reproduce verbatim copies of books or images or whatever, and they try not to do this anyway, and just because you could maybe make an image or whatever with a copyrighted character design or something, normal intellectual property law already restricts you from selling it etc. Seems fine.
cowboylowrez 22 hours ago [-]
Here's the bottom line for me, take two llms, train one on copyrighted materials, do not train the other at all. now, I'm wondering here, which llm will be more sellable, more able to answer questions etc?
How about those music AIs, why aren't they training on classical music and theory textbooks? Why are they being accused of training on copyrighted music? I remember one chat about music AIs, the proponent/enthusiast mentioned "creating" a song by using "johnny cash" in the prompt to get a song in that distinctive style.
I will be the first to admit how useful these AIs are, while I don't use them to "create music", I sure as heck have used them for chats and to write programs. I'm not one of those naysayers talking about the outputs being crap because despite some errors here and there, I've found great utility from these things and I don't even mess with "frontier models" from openai or anthropic. My gripes are about the costs about what we're doing here.
When chatting about "fair use", we can go with the legal definitions which are clear as mud, but I think a better and more sensible path forward would be to consider all the ramifications of their use. Even the term "fair use" implies results much better than we're actually seeing, the economic forces in play here do not seem fair at all.
wilg 13 hours ago [-]
I don't understand your point.
cowboylowrez 13 hours ago [-]
Its my stated disagreement with the following:
>I think it makes sense to say training on material you legally acquired is fair use.
First an aside, there's no law against accessing copyrighted material. OpenAI is certainly welcome to read the New York Times. I can legally buy dvds but the legality of redistributing rips is only considered should I redistribute them.
One of the considerations of "fair use" might be the possibility of benefits or drawbacks to society of those uses that fall under "fair use" exceptions. I would hold the AI's use to that standard, and thats really when we should take the entirety into consideration. This is a big topic tho. We want laws because we believe they benefit society, when loopholes appear we'd normally like them closed, obviously there are branches of the US government that are simply not "normal" right now so theres that lol
I will also admit that a narrower view of "fair use" is to equate "training" of these AIs with human use of the material, after all someone reading the new york times are certainly not infringing on anybody's copyright, in fact they're likely fufilling the new york times internally held purpose, ie., they do the writing, they hope to be read with whatever profit to them that might bring. Just like an AI reading that page by controlling a web browser right? Well this is a good rabbit hole too and you're welcome to try to present that case also, because I've been of the opinion for the last few years that LLMs are not people, and I can chat about that all day.
Its easy to post that you don't understand, so if you'd like to post that again this is fine with me, I'm often a very misunderstood individual :)
wilg 10 hours ago [-]
My position is it's totally coherent to say that "training" is analogous to "reading" and inference is not analogous to "redistributing rips" and that it's a reasonable position for fair use law to operate that way. I think this benefits society.
kova12 21 hours ago [-]
Chinese companies will train their models on all the data they can get their hands on. We can make USA companies act ethically, but then they would become irrelevant. The "safeguards" they were made to implement already makes them less useful than comparable Chinese models. A few more years on that trajectory and we won't have to worry about them anymore
--- The free rider problem is also a form of market failure... The production of public goods results in positive externalities which are not remunerated. If private organizations do not reap all the benefits of a public good which they have produced, their incentives to produce it voluntarily might be insufficient. ---
The US government has just made the NY Times content a public good, and in doing so, allowed openai to create a private good, for which they will be remunerated instead.
It could be argued that this means the US government has effectively taken profit from one company to fund another, under the guise of helping global competitiveness, and whilst this may be the current M.O. of the Trump administration it means predominantly domestic companies are effectively forced into providing export subsidies for the new industry, and monopolising (or oligopolising?) an incumbent.
I wonder if on this basis Seedance could also be equally legally be allowed to use copyrighted media for its video AI, so effectively they're also pulling the legs out from under the film industry too.
I'd much prefer a world where everything is not locked behind a login with bonus 2FA, but that sure seems like where we're headed.
[1] https://news.ycombinator.com/item?id=49491791
NY Times: https://news.ycombinator.com/item?id=49543821
For use as training data you still need some right to it AFAICT.
Which will be also interesting for prompts used by the users of these AI companies, because even if those prompts contain copyrighted data, they can be used for training because training is transforming these prompts into something else.
There's no sense in worrying how closely the university aligns real policy with various theoretical economic ideals.
If we allow copyright law to directly inhibit technological advancement then we defeat the very purpose of copyright law.
That said, it need not inhibit if they just pay for using it. Hopefully they will sue for reasonable compensation, noting that they have such a privilege only in a handful of countries so any restriction is strictly a restriction on those countries, which is a huge technological advantage for all the rest.
I try to see both sides.
You can be right, and dead at the same time.
Just because I think they filed stupid doesn't imply that I think they are wrong.
Not sure if you read my comment in a different light or not, but I do get the feeling you misunderstood my meaning.
I would like clarity on copyright rules and regulations but it's hard to sympathize with NYT when they have been so aggressive with their AI fear-mongering. The authors and artists must be the centerpiece instead of giant corporations. I doubt they will use any leverage to actually improve ethical usage of copyright with AI, and they only want to protect shareholder profits.
https://aiinstitute.hbs.edu/platform-rctom/submission/the-ne...
How about those music AIs, why aren't they training on classical music and theory textbooks? Why are they being accused of training on copyrighted music? I remember one chat about music AIs, the proponent/enthusiast mentioned "creating" a song by using "johnny cash" in the prompt to get a song in that distinctive style.
I will be the first to admit how useful these AIs are, while I don't use them to "create music", I sure as heck have used them for chats and to write programs. I'm not one of those naysayers talking about the outputs being crap because despite some errors here and there, I've found great utility from these things and I don't even mess with "frontier models" from openai or anthropic. My gripes are about the costs about what we're doing here.
When chatting about "fair use", we can go with the legal definitions which are clear as mud, but I think a better and more sensible path forward would be to consider all the ramifications of their use. Even the term "fair use" implies results much better than we're actually seeing, the economic forces in play here do not seem fair at all.
>I think it makes sense to say training on material you legally acquired is fair use.
First an aside, there's no law against accessing copyrighted material. OpenAI is certainly welcome to read the New York Times. I can legally buy dvds but the legality of redistributing rips is only considered should I redistribute them.
One of the considerations of "fair use" might be the possibility of benefits or drawbacks to society of those uses that fall under "fair use" exceptions. I would hold the AI's use to that standard, and thats really when we should take the entirety into consideration. This is a big topic tho. We want laws because we believe they benefit society, when loopholes appear we'd normally like them closed, obviously there are branches of the US government that are simply not "normal" right now so theres that lol
I will also admit that a narrower view of "fair use" is to equate "training" of these AIs with human use of the material, after all someone reading the new york times are certainly not infringing on anybody's copyright, in fact they're likely fufilling the new york times internally held purpose, ie., they do the writing, they hope to be read with whatever profit to them that might bring. Just like an AI reading that page by controlling a web browser right? Well this is a good rabbit hole too and you're welcome to try to present that case also, because I've been of the opinion for the last few years that LLMs are not people, and I can chat about that all day.
Its easy to post that you don't understand, so if you'd like to post that again this is fine with me, I'm often a very misunderstood individual :)