California’s AB 412 Still Demands AI Developers Do The Impossible

from the that's-not-how-any-of-this-works dept

California lawmakers are again considering A.B. 412, a bill that would require AI developers to identify and disclose copyrighted works used to train generative AI systems.

The problem this year is the same as last year: it’s practically impossible to comply with this law. The bill demands information that often does not exist, and cannot realistically be obtained. 

EFF submitted an opposition letter to the California Senate Privacy Committee explaining why we continue to believe A.B. 412 is simply unworkable. To the extent developers do follow this law, it will have the effect of locking in the power of the largest companies in AI. 

A Burden That Can’t Be Met

A.B. 412 sounds simple: just have AI developers create and keep a list of all the registered copyrighted works they use in AI training. 

That may seem straightforward. In practice, it’s anything but. 

There is no machine-readable “list” of copyrighted works at the U.S. Copyright Office. And many copyright holders can get a copyright without even depositing a publicly viewable sample of the work—for example, software companies may register copyright on proprietary code without revealing it to the public. 

And on the open internet, copyright information is often incomplete, unavailable, or impossible to verify. One image may be registered with the copyright office, while the next is licensed under a free Creative Commons license (like the images that EFF creates), and the next is public domain. A message forum user might post an original story, photograph, or poem without any indication of ownership or registration status. 

The bill effectively asks developers to continuously cross-reference massive batches of online data against a copyright system that simply wasn’t designed to do so. If California passes A.B. 412, its impact will go far beyond the large AI companies we read about in the headlines. 

Not Just Big Tech

Supporters often frame this bill as a way to help creative workers have some leverage against Big Tech, but the bill reaches much further than the big AI companies. 

Its definition of “developer” extends to anyone who makes a generative AI model available to Californians. That includes indie developers tinkering with an existing model, open-source initiatives, nonprofits, and other non-commercial efforts. Recent amendments added exemptions for universities and government entities, which is important, but that still leaves out a vast swathe of non-commercial tech work that’s done by people without full-time jobs in government or academia. 

Large companies will hire compliance teams and lawyers to navigate these requirements. Smaller organizations and independent developers usually can’t. The result will be fewer opportunities for startups and new entrants. Faced with this massive compliance burden, some won’t even try. 

Courts Are Already Deciding These Questions

The bill is premised on the idea that copyright owners currently don’t have good remedies if they’re mistreated by AI companies. That simply isn’t true. And the growing wave of federal court filings in this space proves it. Content companies that want to sue tech companies, large or small, have no problem doing so. Those courts are still working through important questions about fair use and transformative use. Some courts have already concluded that many AI training activities qualify as fair use. Others continue to evaluate the issue.

California lawmakers should not rush to impose new state regulation while those questions remain unresolved. This is why copyright is governed at the federal level: both creators and fair users benefit from a single set of nationwide rules. 

At this point, the bill remains a solution in search of a problem. Rights holders already have powerful tools to protect their interests under existing federal law. What this bill adds isn’t clarity or transparency, but a costly and essentially impossible compliance burden that will discourage small developers and researchers. 

California has been able to support both artistic creativity and tech innovation for decades now.  But A.B. 412 does not strike the right balance. 

If you are a California resident and interested in speaking out about this bill, you can find and contact your representatives through this website

Republished from the EFF’s Deeplinks blog.

Filed Under: , , , ,

Rate this comment as insightful
Rate this comment as funny
You have rated this comment as insightful
You have rated this comment as funny
Flag this comment as abusive/trolling/spam
You have flagged this comment
The first word has already been claimed
The last word has already been claimed
Insightful Lightbulb icon Funny Laughing icon Abusive/trolling/spam Flag icon Insightful badge Lightbulb icon Funny badge Laughing icon Comments icon

Comments on “California’s AB 412 Still Demands AI Developers Do The Impossible”

Subscribe: RSS Leave a comment
45 Comments
Bloof (profile) says:

Funny how tech companies can track what advert you looked at on what day, what website, what time, for how many seconds, from what geographical location, using what OS, browser, what plugins you had and how many times you’ve looked at similar things and how many times you’ve clicked through and all those other variables on countless other days… but asking them to make the effort of tracking what material they threw into the slop machine, that’s beyond their ability! That is too much of a burden, you see. Stifling innovation, yadda, yadda.

Won’t someone please think of the mom and pop AI startups that may appear at some unspecified point in the future?

MrWilson (profile) says:

Re: Re: Re:

That means only large corporations and venture-capital-flush startups would be able to afford to develop the systems that are going to dictate your software experience with your insurance, health care, social services, consumer experiences, etc.

You’d need to do away with capitalism for that stance to mean anything useful.

Bloof (profile) says:

Re: Re: Re:2

You say that like the bulk of what is happening right now isn’t being done by big tech and startups, flush with VC cash bent on finding new ways to track people or make them unemployed.

It isn’t mom and pop startups or bedroom coders being given government AI contracts or developing those systems, it’s open AI, Palantir and so on, stop using theoretical little guys as a figleaf.

MrWilson (profile) says:

Re: Re: Re:3

Except they’re not theoretical little guys. They do exist and they won’t be able to get any contracts if they’re required to pay big money to train their models to any level of effectiveness. You’re functionally saying that because billionaires can afford to pay for licensing, they’re the only ones who should be able to develop LLMs or get contracts. This stance only works if you’re just in for “all LLMs are bad for everything and shouldn’t exist” and you have an effective plan for wiping them out across the world.

If my health care provider decides that my medical care is going to be dictated by an LLM, I want it trained on the breadth of medical journals and research, not Grok’s eugenicist white supremacist bullshit. Musk can afford to license that training material while ethical developers wouldn’t be able to.

Epic_Null (profile) says:

Re: Re: Re:4

Honestly, I would like to challenge your premises here.

Any company that wants to stay arround has to charge at least the cost of doing buisness. These LLMs, on top of being heavily subsidized, have not had to take into account the cost of licencing. If we force big companies to pay licencing costs, they won’t be able to continue selling LLMs.

Speaking of licencing costs: this only matters if copyright holders say “Yes”. If they say “No”, as they have a legal right to do so, then LLM companies start having an input problem.

Licencing costs also don’t just matter during training. It’s not at all unusual for licences to include royalties, meaning licence fees need to be paid out for the lifetime of the model, not just the training period.

SO lets say we started enforcing the law here:

  • Enfringing models would have to be retired, since many of the copyright holders would refuse to authorize use of their content.
  • LLMs get even more expensive, and companies do the math. They replace the LLMs with humans and things can chill some.
  • Ethical Developers no longer need to compete with GROK, since GROK’s enfringement has caused it to be retired by legal means
  • The list of copywritten works allows work authors to skip proving their work was used in trainig, making their cases easier to build.
MrWilson (profile) says:

Re: Re: Re:5

Any company that wants to stay arround has to charge at least the cost of doing buisness.

Not necessarily. Open source developers, non-profits, and independent researchers aren’t “companies.” There are operational costs, but sometimes those are covered by the university they work for, through grant funding, etc.

If we force big companies to pay licencing costs, they won’t be able to continue selling LLMs.

Of course they will. They’ll just make deals with large publishing platforms that will pay the big authors a lot of money and smaller authors pennies or nothing, similar to music streaming services or ASCAP. They just don’t want to because it means less profits.

The other part is that they can operate in a gray area where if they don’t disclose the source of their training data, you’d have to find your own evidence to prove that they used your works. Since the training data doesn’t exist in the models themselves, you’d need a disclosure of some kind.

Speaking of licencing costs: this only matters if copyright holders say “Yes”. If they say “No”, as they have a legal right to do so, then LLM companies start having an input problem.

Many publishers reserve rights to make licensing decisions for authors, so it won’t necessarily be the millions of authors. It’ll be Jeff Bezos saying, “sure, pay me and it’s yours.” There might be fallout, but Amazon has a bulk of the publish market already and is already boycotted by people who don’t like them and it hasn’t hurt their bottom line much.

Licencing costs also don’t just matter during training. It’s not at all unusual for licences to include royalties, meaning licence fees need to be paid out for the lifetime of the model, not just the training period.

No, that wouldn’t be something the LLMs would agree to and if the publishers or the authors want to get paid, they’ll have to compromise. There will be enough copyright holders who will be fine with it, so the LLM companies can just exclude the ones who have higher expectations. They might do it for a big name author, but they’re not paying royalties to an author who sold ten copies of a book to his friends and family members. They want to ingest as much as possible for better models, so the price point per each work or author will be minimal and there will be enough willing to get paid anything at all that there won’t be as much of a bargaining position. The big LLM companies are opportunistic tech bro hustler assholes. They’re not going to agree to terms that are too expensive for them. They’ll treat 99.99% of authors the same way they do gig economy workers.

SO lets say we started enforcing the law here:

While there’s an argument to be made for criminal copyright infringement, it’s more commonly a civil lawsuit, so we wouldn’t start enforcing the law, the affected plaintiffs would have to lawyer up and sue, which means they’re already out money unless the lawyers see enough certainty and dollar signs.

Enfringing models would have to be retired, since many of the copyright holders would refuse to authorize use of their content.

You’d have to prove which works were used first. You’d have to subpoena and get the training data, which may already be gone in some scenarios. And the companies know how to train models, they’ll just train them without the plaintiff’s work.

LLMs get even more expensive, and companies do the math. They replace the LLMs with humans and things can chill some.

Except the companies pushing LLMs don’t want the humans. They’re not looking for humans as a solution. They embrace LLMs because they perceive that humans are more expensive individually in the long term and more problematic. They want max profits. They don’t care about environmental impacts or climate change. Everything is an opportunity and every employee is a financial expenditure and liability. If you ban LLMs, they’ll just rename them “advance database scripts” and keep going until the next lawsuit they have to pay off.

Ethical Developers no longer need to compete with GROK, since GROK’s enfringement has caused it to be retired by legal means

This bill is only in California. If it becomes law and isn’t shot down in the courts (and it will most definitely be challenged by big money), developers will just move out of California. And this only hurts US firms. China won’t stop its development. Foreign developers can violate copyright all they want as long as their country’s government isn’t overly deferential to the US (and who is to say if Trump would worry about it unless he could just extract a bribe to let it go…).

The list of copywritten works allows work authors to skip proving their work was used in trainig, making their cases easier to build.

That’s ignoring how difficult it is to attribute works, which is covered in the article.

Try attributing my comments. You don’t know my full name. Maybe I’m quoting someone else and the copyright belongs to them. Maybe I’m repeating myself and the copyrighted content has a different date than you thought. What’s the name of the work? How would you find it in a list? List? You’d need a database with a really good search function. Who is hosting the database? Who is paying the costs of hosting the database? Is the developer or the state of California? Do I have to pay for a subscription to access it like a journal or a court document system?

I understand the desire to hold big corporations accountable, but we need societal and cultural changes, including changes to capitalism itself. LLMs pirating works is just one symptom. Wealthy corporations game the system such that they also game the courts and game even the few laws that seek to actually restrain them. And kneecapping open source developers by creating a cost of doing business that the wealthy corporations will definitely be able to afford is just giving the big LLM companies more of a leg up than they already have. They’ve asked in some cases to be regulated so that they can engage in regulatory capture and keep our competition.

Bloof (profile) says:

Re: Re: Re:4

Hate to break it to you, but if your Healthcare provider is using ab AI model to dictate your care, they aren’t picking the expensive one trained legally on the finest medical research, they’re picking the one that will say no, and that’s more likely to be OpenAI based if you’re lucky, Grok if you’re not.

MrWilson (profile) says:

Re: Re: Re:5

They claim they’re not using LLMs for diagnoses or treatment plans yet, but my health care provider uses Abridge for documentation (which can affect care determinations) which is based on an Nvidia Nemotron open model, but they use OpenAI GPT models as well. They also claim to use the H20.AI agentic LLM platform for analytics and forecasting of patient trends.

This comment has been deemed insightful by the community.
Anonymous Coward says:

Re:

Funny how tech companies can track [list of things that tech companies decide to track and then DO track], but asking them to make the effort of tracking [what they’ve already done without tracking], that’s beyond their ability!

So… you can track for me what kind of sandwiches you’ve had for lunch, going back the last five years, right? I mean, it’s easy to track, right?

This comment has been deemed insightful by the community.
Anonymous Coward says:

Re:

“It was the best of times, it was the worst of times, it was the age of wisdom, it was the age of foolishness, it was the epoch of belief, it was the epoch of incredulity, it was the season of Light, it was the season of Darkness, it was the spring of hope, it was the winter of despair, we had everything before us, we had nothing before us, we were all going to Heaven, we were all going to Hell.”

So tell me…
* Was that from the novel? (Public Domain)
* The 1958 movie? (copyrighted)
* A transcript from the 1938 radio play? (copyrighted by someone else)
* From the film “The Only Way”? (copyrighted if ingested prior to 2023, public domain if ingested after)
* From “some unidentified source from 1927”? (copyrighted if a sound recording, public domain if a “published source”)

Remember, if you get it wrong, the lawyers suing you win.

Bloof (profile) says:

Re: Re:

Ai companies pretending they either don’t track or can’t track the origins of their training material is not the gotchya you think it is, champ. If they can’t identify where something came from and whether or not they have any legal right to use it, they shouldn’t use it.

The world is going to sh-t because of entitled assholes doing crime in bulk and pretending that makes it legal and right, and it’s unfair to them to obey laws that say they should have to obey the laws that were in place when they did the thing in the first place.

Rocky (profile) says:

Re: Re: Re:

Ai companies pretending they either don’t track or can’t track the origins of their training material is not the gotchya you think it is

Sure, but it wasn’t really about tracking origins but copyrighted works. Determining if something is copyrighted or not can be very difficult regardless of the origin which some people find to be a very good feature for monetization. Aside from that, just ingesting any kind of content without due diligence is just plain greedy and stupid and those who do that can just fuck off.

This comment has been deemed insightful by the community.
Epic_Null (profile) says:

The bill demands information that often does not exist, and cannot realistically be obtained.

Why can’t it be obtained, Joe? WHY CAN’T IT BE OBTAINED?!

Look. I work with software. You give me any, ANY of our codebases and half an hour, and I will get you at least a shallow dependency list. Give me a day and I can script out some shit to process that into something actually usable.

You don’t just shove things in your software and call it a day, developers HAVE A RESPONSIBILITY to manage the Other People’s Work that is included in the final product.

And it’s not just us either. I bet if I walked into a resturaunt and asked the owner “Who are your suppliers, you have one day to answer”, they could get me a list of every supplier for a year or point me to the people who could. Hell, there are buisnesses where I could walk in with a product serial number and ask about the source of a specific washer used in the product, and they would be able to point me to the specific LOT that the washer came from.

So when I say I have ZERO sympathy for how hard it is to produce the list of worke that they used to train their products, I absolutely mean it.

So what if you don’t know “Is this copywritten”?! If you didn’t know then:

  • you should probably include it in your list
  • why the hell are you using it if you don’t know its legal status
  • No seriously what the fuck were you thinking including a dependency you did not bother to understand the legal status of?

There is, what I will call, a minumum standard for doing buisness. If you don’t meet that standard, you should not be doing buisness. Being able to answer “How do you source your inputs” is part of that minimum standard.

terribly tired (profile) says:

Re:

Hell, there are buisnesses where I could walk in with a product serial number[…]

Yep. Anyone who manufactures anything remotely complex at anything remotely resembling scale needs this sort of granularity.

Even if you yourself don’t ever need any of the information, because your enginerding/proto, production, QA, logistics, and aftermarket depts are all absolutely flawless you might still be contractually bound to be able to provide exactly this sort of info to any number of customers upon request.

Contracts and entire relationships are lost over much, much less.

Anonymous Coward says:

Having a working brain, or, at least, a few working brain cells, has never been a requirement for election to public office. This bill is ample proof! The net effect would be that no AI developers/companies would have a presence in California. All those empty datacenters could be made into living space for the homeless…

Anonymous Coward says:

Of COURSE this can be done

I run an archive which has over 10 million items in it at the moment and will likely hit 25 million by 2030. I can tell you the provenance and copyright status of every one of them, because — unlike the AI companies — I’ve done everything possible to respect the rights of the people who created those materials, copyright and otherwise. Most of that was accomplished by doing a strange, wonderful, magical thing that you might have heard of: it’s called asking permission. I know, I know, it’s exotic and difficult to grasp, but there you have it.

Any of these AI companies could solve this problem almost trivially by putting 1% of their funding into it. If they did so, they would probably be spending thousands of times more than I’m spending doing it — because we’re a nonprofit and can’t spend that much. So, with that much money, they should be able to achieve vastly superior results in much shorter time. But they haven’t and they’re just whining about it either directly or through their well-paid shills — like you.

Your entire premise is absolute bullshit. These companies could easily do this, it’s just that they don’t want to. They want to steal, steal, steal.

Anonymous Coward says:

Re:

I have to ask where you find time to read all of the responses to your requests. I mean, 10 million? Whew! If you took 10 seconds per response to read the response, associate it with the item you asked permission for, and hit the “go” button, that’s something like 190 calendar years of work. Where do you find the time?

Strawb (profile) says:

Re:

I could understand the skepticism if this was written by a big tech executive, but it’s coming from an organization that tends to understand the intersection of tech and law.

And as he points out, the issue isn’t that the big companies probably couldn’t throw money at the problem to fix it(or get around it); it’s that the small AI developers don’t have the money to throw at the problem. So it would further entrench AI development powers for Big Tech.

This comment has been deemed insightful by the community.
Anonymous Coward says:

Nah, fingerprinting and listing all the input material is a trivial task for any developer who has passed compsci 101. There’s nothing hard about it.

The only part of the bill that might be problematic for a developer to comply with is “3115.5 (a) (3) Document the rights owner of each covered material documented pursuant to this subdivision.” And since they are only required to make reasonable efforts to identify and document material, this is hardly a meaningful gotcha.

The bill effectively asks developers to continuously cross-reference massive batches of online data against a copyright system that simply wasn’t designed to do so.

It does not. Including all material used (registered or not) in the disclosure is perfectly compliant with this law. A developer may choose to cross-reference if they wish to keep more of their material secret, but that is a choice they made, not a requirement of the law.

At worst, a cottage industry would pop up to produce and sell a mirror of the copyright system which is designed fo this. Which, frankly speaking, is actually the best case scenario as it opens that data to the public in a way which the government currently refuses to do.

This comment has been deemed insightful by the community.
Arianity (profile) says:

it’s practically impossible to comply with this law. The bill demands information that often does not exist, and cannot realistically be obtained.

Not really sure why that’s our problem that their model of indiscriminately scraping isn’t viable. Something that anyone wanting to use large amounts of copyrighted works would have to deal with as well.

Large companies will hire compliance teams and lawyers to navigate these requirements.

You just said it wasn’t possible. If it’s actually impossible, it’s not something you can simply throw compliance teams and lawyers at.

Smaller organizations and independent developers usually can’t.

Content companies that want to sue tech companies, large or small, have no problem doing so.

The U.S. legal system is famously expensive and unavailable for smaller litigants, a point you yourself are literally making in this article, and something EFF is well aware of. It’s also not just content companies who have copyright protections- it’s smaller organizations and independent creators, too.

Rights holders already have powerful tools to protect their interests under existing federal law.

And yet, you couldn’t give a single resolved example actually doing so. Never mind that EFF likely wouldn’t support any such decision if it were actually resolved in favor of rights holders in the first place.

Boba Fatt (profile) says:

EVERYTHING is copyrighted

Under US law and most other countries (Berne Convention) nearly every work gets a copyright as soon as it’s created. The copyright owner – which is not necessarily the author – has the right to control who can make copies, display the work, make derivative works, etc.

Most of the outraged replies saying the EFF is wrong include quotes from the first post. THAT QUOTED TEXT IS COPYRIGHTED. By posting those quotes, you’re using someone else’s copyrighted content.

The US has fair use exceptions, but many other countries don’t, and even in the US that’s only useful as a defense in a lawsuit – it’s not a right in itself.

Who owns the copyright on those quotations? The EFF? Techdirt? Someone else? Has the owner of the EFF’s statements granted permission for Joe Mullin or Techdirt to post them in the article? Can Techdirt grant permission for commenters here to quote it further? Have I granted permission for THIS post to be quoted? That’s the problem.

Sure, automate that fact-finding process for every single thing on the web. I’ll wait.

Yes, AI scrapers are jerks and they’re abusing long-standing social conventions in ways that are actively harmful. Techdirt has posted multiple articles about that. But now it’s bad enough to trigger the politician’s syllogism, which makes it even worse. That’s what the EFF is opposing.

Anonymous Coward says:

Re:

“Most of the outraged replies saying the EFF is wrong include quotes from the first post. THAT QUOTED TEXT IS COPYRIGHTED. By posting those quotes, you’re using someone else’s copyrighted content.”

Of course we/they are. And of course this is fair use under copyright law, because we’re merely quoting a very small portion of it for reference. You know this full well, you’re just pretending that you don’t.

“Sure, automate that fact-finding process for every single thing on the web. I’ll wait.”

While I’m certain that this is well beyond your feeble abilities, it’s not beyond those of others: it’s just a matter of having enough money to fund the resources required to do it. And every single one of these AI companies has enough money to do this thousands of times over — and they certainly have enough computing capacity and staff to do it.

In other words, you’re lying about the difficulty of the problem in order to make it seem intractable. It’s not.

Epic_Null (profile) says:

Re:

Thing is, for something like a comment on the article, everything is “Reasonably Cited” – by which I mean if I quoted your quote, and someone had just the RSS feed entry of my quoting your quote, the author would be quickly discovered.

The law as described in the article is only asking companies to “identify and disclose copyrighted works”, which could easily be done in the case of quoting a quote of a quote.

You would simply cite where you got it from and pass along relevant citations from where you got it from, referencing all relevant authors as impacted copyright holders. After all, the message containing the quote is a derivative work.

Sure, automate that fact-finding process for every single thing on the web. I’ll wait.

Well… it’s probably doable anywhere that consented to be used for AI training. But you’re right about it not being practical everywhere. Which IMO is an argument that said content should not be used in the first place. Our laws should not be written to protect “jerks […] abusing long-standing social conventions in ways that are actively harmful” from the consequences of their failure to have common decency.

This comment has been deemed insightful by the community.
Anonymous Coward says:

Re:

If you’re not considering edge cases, you’re not answering the problems with a law that does not recognize edge cases.

So, just for instance: If the work was under copyright when it was ingested, but is public domain now, what happens?

Anonymous Coward says:

Re: Re:

If they are unsure of the copyright status, they can simply include it in the dataset. The law does not require that the dataset only includes registered copyrighted material, it requires that all such material is in the dataset. Including any material they are unsure of, or just including all material they used without even attempting to check the status, is perfectly compliant.

Anonymous Coward says:

i don’t care to get into how bad copyright law is (it’s terrible), or how bad stupid legislation is. One must think around the popularframing of the problem. The answer is, “AI devs can FOAD”. i don’t care if a fistful of ppl use it properly and find it useful. What it is, is an amplifier of everything wrong with business and people. All tge things that are (properly!) analyzed at this blog for over 20 years as being not-good – that’s what current, commercial AI is for. It was built for the purpose of enshittification and value extraction, and that’s how it is used. Even if some use it so poorly it is also to their own detriment.

Add Your Comment

Your email address will not be published. Required fields are marked *

Have a Techdirt Account? Sign in now. Want one? Register here

Comment Options:

Make this the or (get credits or sign in to see balance) what's this?

What's this?

Techdirt community members with Techdirt Credits can spotlight a comment as either the "First Word" or "Last Word" on a particular comment thread. Credits can be purchased at the Techdirt Insider Shop »

Follow Techdirt

Techdirt Daily Newsletter

Subscribe to Our Newsletter

Get all our posts in your inbox with the Techdirt Daily Newsletter!

We don’t spam. Read our privacy policy for more info.

Ctrl-Alt-Speech

A weekly news podcast from
Mike Masnick & Ben Whitelaw

Subscribe now to Ctrl-Alt-Speech »
Techdirt Deals
Techdirt Insider Discord
The latest chatter on the Techdirt Insider Discord channel...
Loading...