Just last week we had the story of the Trump DOJ issuing very questionable subpoenas of NY Times journalists while trying to track down who leaked information to those reporters regarding the potentially catastrophic security flaws of the “gift” 747 plane he received from Qatar. As we noted in that original post, this appeared to be a wholly abusive use of the government’s subpoena powers, and well outside the norm.
On Thursday, the DOJ agreed to withdraw those subpoenas, but only after a long court hearing in which the DOJ thoroughly embarrassed itself in front of the judge, Arun Subramanian, who noted many, many problems with the subpoenas, which the DOJ tried to tiptoe around, calling them “inadvertent errors.” Most of the media coverage of this is pretty weak, but Matthew Russell Lee of the Inner City Press did a wonderful liveposting of the hearing that suggests just how badly the DOJ fucked this up.
It started out with the DOJ saying they weren’t going to withdraw the subpoenas, and claiming that they believed the subpoenas were “properly” issued. But the court quickly pointed out that there is precedent in the Second Circuit regarding when and how you can subpoena journalists, and the DOJ basically ignored all of that. The DOJ’s Sean Buckley argued that following those rules would amount to conceding the rules applied — something the DOJ apparently didn’t want to admit, leading the judge to say that following the rules wouldn’t be seen as any such admission.
Buckley: If we withdraw it might imply we accept the Gonzalez test.Judge: You're going to go to the Supreme Court?Buckley: It's possible. Judge: I will not understand any withdraw as accepting the Gonzalez test. I'm trying to figure out a practical way here
Judge Subramanian kept pressing Buckley on why the DOJ rushed to issue these incredibly broad subpoenas when there appeared to be much more straightforward ways to obtain the information they were seeking. Indeed, another part of what was discussed is that the DOJ’s subpoenas were so broad that they included phone records of reporters’ relatives who had nothing whatsoever to do with the reporting:
A Justice Department lawyer, Sean Buckley, cast the government’s missteps as inadvertent errors and said: “No one was trying to pull a fast one.” Buckley apologized for other subpoenas that sought records for phone numbers belonging to one reporter’s mother and two of the journalists’ spouses.
“That was an error, judge, which we own,” Buckley said. “It was a consequence of trying to move quickly.”
“These things are starting to pile up,” Subramanian said, becoming increasingly testy.
The judge also explored whether or not the DOJ misled the judge who signed off on the subpoenas, by not letting them know that the subpoenas were for information associated with reporting. He even noted that the Assistant US Attorney who got the subpoenas, Kevin Sullivan, was in the room, but not at the table, asking him to come out of the galley and join the DOJ table (this is not something that usually happens).
Buckley: We would be prepared to immunize these reporters – our focus is on the leakers.Judge: Is Mr. Sullivan here?Buckley: Yes. In the galleryJudge: Why? He is on all the pleadings. Come up- we have extra chairs here. Good afternoon.Sullivan: Good afternoon
Following that was an incredible exchange wherein Judge Subramanian asked Sullivan about whether he told the original subpoena-issuing judge that the subpoenas were for reporters, leading Sullivan to say it “was an oversight” and that later on they “did legal research.”
Judge: You didn't tell the judge that the subpoena was about reporters, about the New York Times?Sullivan: We did not. It was an oversight. Later we did legal research.Judge: Wouldn't it have been relevant to know there had been public reporting? A: Yes
Around that point, a clearly fed up Subramanian said that if this were a normal case, this would be the point where he would issue an order to show cause why the DOJ shouldn’t face sanctions for abusing the subpoena process. There was some more back and forth scolding, including Subramanian pointing out that the “errors” for the DOJ seemed to be “piling up” and asking the DOJ if he should expect to see more mistakes like this moving forward.
Around this point, the DOJ regrouped and changed their stance from earlier in the hearing, saying they were now willing to drop the subpoenas. After the hearing was over, the ever petulant Trump Justice Department quickly whined to the media how unfair it was that federal judges expected them to actually follow the rules and stuff:
After the hearing, the Justice Department lashed out at Subramanian in a statement, saying he “threatened our attorneys with sanctions unless subpoenas were withdrawn, and blocked us from presenting the meticulous process of this investigation.”
“The grand jury has a right to hear testimony from all material witnesses in a federal criminal investigation. This judge’s conduct overrides clear longstanding principles and common sense — blocking the grand jury from receiving core evidence in a national security investigation,” the statement said.
“Make no mistake,” it added, “this investigation remains ongoing, and we will pursue justice against those threatening national security by leaking classified information, a serious federal crime.”
Once again, we have an overly aggressive, understaffed, and generally incompetent DOJ that seems not to realize that there are significant and important Constitutional limits on what it can do. And when a judge calls them out on it, the fact that their immediate response is to start whining about it like they were the victims here suggests a good reason that the entire DOJ will need a massive overhaul post-Trump.
Since the New York Times published its semi-viral big profile of Medvi last week — the “AI-powered” telehealth startup that it breathlessly described as a “$1.8 billion company” supposedly run by just two brothers — I’ve had multiple friends and family members send me the article with some version of the same message: “Can you believe this guy built a billion-dollar company with AI? Why haven’t you done this?” The story is making rounds, and giving people the impression that with a ChatGPT account and a little bit of marketing know-how, you too could be raking in millions every month.
The problem is that most of the story is utter nonsense.
Let’s start with the headline number itself. The NYT admits — buried deep in the piece — that Medvi “has not raised outside funding” and “has no official valuation.” A company’s value is typically established by investors, an acquisition offer, or public market pricing. Medvi has none of those. What it has is a revenue run rate — a projection based on early-2026 sales extrapolated across a full year. Calling that a “$1.8 billion company” is like calling someone who found a twenty on the sidewalk a “future millionaire.” Any business reporter should know the difference. Even the NYT tips its hand:
Medvi is technically not a one-person $1 billion company, since Mr. Gallagher hired his brother and has some contractors. The start-up, which has not raised outside funding, also has no official valuation.
“Technically not” doing quite a bit of heavy lifting there.
But the misleading valuation is almost the least of it. Even if you accept revenue as the relevant metric, how sustainable is that run rate for a company that just got an FDA warning letter, is facing a class action lawsuit for spam, has a key partner being sued over allegations that a major product doesn’t actually work, and is operating in an industry that regulators are actively trying to rein in?
Oh, wait, did the NYT forget to mention all of those things? They sure did! Not to mention the legions of fake, apparently AI generated doctors and patients who keep showing up in Medvi advertisements. Yes, the NYT eventually alludes to some of that, but it claims these were mere “shortcuts” that were fixed last year (they weren’t).
That said, you can feel the pull of the narrative that seduced the NYT: a scrappy founder with a rags-to-riches backstory, two brothers taking on the world, AI tools stitching it all together, Sam Altman himself anointing the achievement as proof that his prediction of a “one man, one billion dollar company, thanks to AI” was correct.
It’s a hell of a story. The problem is that almost none of it holds up to even the most basic scrutiny, and the fact that the New York Times — the New York Times — fell for it (or worse, didn’t care) is an embarrassment. As much as I’ve made fun of the NYT for its bad reporting over the years, this is (by far) the worst I’ve seen. They didn’t just misunderstand something, or try to push a misleading narrative, they got fully played on a bullshit story that any competent reporter or editor should have realized from the jump. This one stinks from top to bottom.
Medvi’s success has very little to do with “AI” and quite a lot to do with fake doctors, deepfaked before-and-after photos, misleading ads, probable snake oil, and the kind of old-fashioned deceptive marketing that has been separating marks from their money for centuries. The only thing AI really “turbocharged” here was the company’s ability to generate bullshit at scale. Oh, and also the NYT somehow missed out on the FDA already investigating the company, as well as the multiple lawsuits accusing the company and its partners of extraordinarily bad behavior.
Let’s start with what the NYT actually published. Reporter Erin Griffith’s piece reads like a press release that the NYT re-formatted as a newspaper article:
Matthew Gallagher took just two months, $20,000 and more than a dozen artificial intelligence tools to get his start-up off the ground.
From his house in Los Angeles, Mr. Gallagher, 41, used A.I. to write the code for the software that powers his company, produce the website copy, generate the images and videos for ads and handle customer service. He created A.I. systems to analyze his business’s performance. And he outsourced the other stuff he couldn’t do himself.
His start-up, Medvi, a telehealth provider of GLP-1 weight-loss drugs, got 300 customers in its first month. In its second month, it gained 1,000 more. In 2025, Medvi’s first full year in business, the company generated $401 million in sales.
Mr. Gallagher then hired his only employee, his younger brother, Elliot. This year, they are on track to do $1.8 billion in sales.
A $1.8 billion company with just two employees? In the age of A.I., it’s increasingly possible.
And then, because no AI hype piece would be complete without the requisite papal blessing from San Francisco:
In an email, Mr. Altman said that it appeared he had won a bet with his tech C.E.O. friends over when such a company would appear, and that he “would like to meet the guy” who had done it.
Altman “would like to meet the guy.” Well of course he would! The NYT hand-delivered him the perfect anecdote for his next AI hype session. The reporter seemingly solicited that quote to validate a pre-existing thesis: “Sam Altman was right about one-person billion-dollar AI companies.” The fact that the company is a dumpster fire of regulatory violations and consumer fraud was, apparently, a secondary concern to the “Great Man and A Great AI” narrative of innovation. This piece was built around a thesis — Sam Altman was right — and then a company was located to prove it.
To its minimal credit, the NYT does kind of acknowledge — eventually, if you make it past the thirtieth paragraph — that things weren’t entirely on the up and up:
Medvi’s initial website featured photos of smiling models who looked AI-generated and before-and-after weight-loss photos from around the web with the faces changed. Some of its ads were AI slop. A scrolling ticker of mainstream media logos made it look as if Medvi had been featured in Bloomberg and The Times when it had merely advertised there.
I mean… shouldn’t that have raised at least one or two red flags within the NYT offices? Medvi’s website featured a scrolling ticker of media logos — including the New York Times logo — to make it look like these outlets had written about the company, when they hadn’t. A year ago, Futurism’s Maggie Harrison Dupré had even called this out directly (along with Medvi’s penchant for bullshit AI slop advertising).
Just underneath these images, MEDVi includes a rotating list of logos belonging to websites and news publishers, ranging from health hubs like Healthline to reputable publications like The New York Times, Bloomberg, and Forbes, among others — suggesting that MEDVi is reputable enough to have been covered by mainstream publications.
…. But… there was no sign of MEDVi coverage in the New York Times, Bloomberg, or the other outlets it mentioned.
And then, despite this, the New York Times went ahead and wrote the glowing profile that Medvi had been falsely claiming existed. The paper of record became the validation that the fake credibility ticker was trying to manufacture.
And the NYT frames all of what most people would consider to be “fraud” as mere “shortcuts” that the founder later “fixed.” Eighteen paragraphs after burying the admission, it reports:
That gave Matthew Gallagher breathing room to fix some shortcuts he had initially taken, like swapping out the before-and-after weight-loss photos for ones from real customers.
“Shortcuts.” Using deepfake technology to steal strangers’ weight-loss photos from across the internet, alter their faces with AI, give them fake names and fabricated health outcomes, and pass them off as your own satisfied customers — that’s a “shortcut.” Ctrl-F is a shortcut. This sounds more like fraud.
And it turns out those “shortcuts” hadn’t actually been fixed at all. As Futurism’s Dupré reported in a follow-up piece published after the NYT article:
As recently as last month, nearly a year after the NYT said that Medvi had cleaned up its act, an archived version of Medvi.org shows that it was again displaying before-and-after transformations of alleged customers. They bore the same names as before — “Melissa C,” “Sandra K,” and “Michael P” — and again listed how many pounds each person had purportedly lost and the related health improvements they apparently enjoyed.
Even though they had the same names, these people that the site now called “Medvi patients” now looked completely different from the original roundup of Melissas, Sandras, and Michaels. Worse, some of the images now bore clear signs of AI-generation: the new Sandra’s fingers, for example, are melted into her smartphone in one of her mirror selfies.
They kept the same fake names and the same fake weight-loss numbers but swapped in entirely different fake people. What the NYT claims was “fixing shortcuts” appears to actually be just “updating the con.”
In a great takedown video by Voidzilla, it’s revealed that at least one set of original images appeared to have been sourced from Reddit forums on weight loss having nothing to do with Medvi, and even with the modified images it used, it massively overstated how much weight the original person claimed to have lost. And while Medvi later switched out the photos with someone totally different, they kept the same name and same false weight loss claims.
And again, all of this was publicly known information that Griffin or her editors could have easily found with some basic journalism skills. We already mentioned that Futurism article from May of 2025, nearly a full year before the NYT piece ran. That investigation traced the deepfaked before-and-after photos back to their real sources, found that a doctor listed on Medvi’s site had no association with the company and demanded to be removed, and documented the AI-slop advertising. That investigation was widely available. A Google search would have found it.
But the fake photos and fraudulent branding are almost quaint compared to what the NYT chose not to mention at all. Six weeks before the NYT piece was published, the FDA sent Medvi a warning letter for misbranding its compounded drugs. The letter admonished Medvi for marketing its products in ways that falsely implied they were FDA-approved and for putting the “MEDVI” name on vial images in a way that suggested the company was the actual drug compounder. The letter warned:
Failure to adequately address any violations may result in legal action without further notice, including, without limitation, seizure and injunction.
The NYT did not mention this letter. And yes, Gallagher now insists that the FDA letter was targeting an affiliate that was using a nearly identical name, and it was that rogue affiliate that was the problem. But the letter is addressed to MEDVi LLC dba MEDVi, which is the name of his company. If he’s allowing affiliates to use his exact name, then that alone seems like a problem. Indeed, it certainly seems to highlight how this is all just, at best, a pyramid scheme of snake oil salesmen, where Gallagher has affiliates willing to deceive to sell more snake oil.
Separately, on March 20, 2026 — thirteen days before the NYT piece ran — a class action lawsuit was filed against Medvi in the Central District of California alleging that the company uses affiliate marketers to blast out deceptive spam emails with spoofed domains and falsified headers. The complaint alleges Medvi is responsible for over 100,000 spam emails per year to class members. The lawsuit seeks $1,000 per violating email.
The NYT did not mention this lawsuit either, even as it was yet another bit of evidence that either Medvi is up to bad shit, or it has a bunch of out of control affiliates potentially breaking laws left and right to increase sales.
A Drug Discovery & Development review conducted on April 3 of MEDVi’s website, Facebook advertising and public records found a pattern of apparent AI-generated personas, including some presented with medical titles, alongside marketing practices that appeared to go beyond the issues identified so far by regulators. A search of Meta’s Ad Library for “medvi” returned more than 5,000 active ads, many of them running under fabricated physician personas. One Facebook page for “Dr. Robert Whitworth,” which ran sponsored ads for MEDVi’s QUAD erectile dysfunction product, was categorized as an “Entertainment website” and listed an address of “2015 Nutter Street, Cameron, MT, 64429,” a location that does not appear to exist. Other ads ran under names including “Professor Albust Dongledore” and “Dr. Richard Hörzgock,” used AI-generated video testimonials and recycled identical scripts across multiple fabricated personas. In several cases, the page displayed a doctor headshot while the ad itself featured an unrelated person delivering a patient testimonial.
After public scrutiny following the article, those fake doctor accounts started disappearing. In fact, Medvi’s own website fine print acknowledges the practice:
Individuals appearing in advertisements may be actors or AI portraying doctors and are not licensed medical professionals.
Seems like maybe something the NYT should have noticed?
Oh, and that same Drug Discovery and Development article highlights how other snake oil sales sites are using the same named doctors… but with totally different images.
Same names… different people. Drug Discovery and Development has a bit more info about Drs. Carr and Tenbrink:
MEDVi’s current site lists two physicians: Dr. Ana Lisa Carr and Dr. Kelly Tenbrink. Both are licensed doctors who work together at Ringside Health, a concierge practice in Wellington, Florida, that serves the equestrian community. Neither is identified on MEDVi’s site as being affiliated with Ringside Health. On MEDVi’s site, Dr. Tenbrink is listed under “American Board of Emergency Medicine.” Dr. Carr is listed under St. George’s University, School of Medicine, her medical school. The Florida Department of Health practitioner profiles for both physicians state that neither “hold any certifications from specialty boards recognized by the Florida board.” A search of the American Board of Emergency Medicine‘s public directory, which lists 48,863 certified members, returned no current affiliation for Dr. Tenbrink.
Did the NYT do any investigation at all? Serving the equestrian community?
Even the few real doctors Medvi claims to work with turn out to be questionable. From Futurism’s article from last May (again, something the NYT should have maybe checked on?):
We contacted each doctor to ask if they could confirm their involvement with MEDVi and NuHuman. We heard back from one of those medical professionals at the time of publishing, an osteopathic medicine practitioner named Tzvi Doron, who insisted that he had nothing to do with either company and “[needs] to have them remove me from their sites.”
Then there’s what a class action lawsuit filed last November against Medvi’s main partner, OpenLoop Health, alleges about the actual products being sold. The NYT frames OpenLoop as basically making what Gallagher is doing possible, noting that while Gallagher has his AI bots creating marketing copy OpenLoop handles: “doctors, pharmacies, shipping and compliance.” You know, the actual business.
So it seems kinda notable that way back in November of last year, this lawsuit was filed that claims that the compounded oral tirzepatide tablets — one of Medvi’s key offerings — are essentially pharmacologically inert when delivered as a pill. Tirzepatide (marketed as Zepbound by Eli Lilly) is an FDA approved weight-loss drug as an injectable. But OpenLoop and Medvi have apparently been selling it in pill form. And Eli Lilly says that there are no human studies, let alone clinical trials, involving any tirzepatide pills.
All of that seems like the kind of thing reporters from the NYT should point out.
What we actually have here is a marketing operation that used AI to automate the production of deceptive advertising at a scale and speed that would have been harder to achieve otherwise. Snake oil salesmen have existed forever. What AI gave Matthew Gallagher (and, I guess, his affiliates) was the ability to crank out fake doctors, fabricated testimonials, and deepfaked before-and-after photos faster than any human team could — and to do it cheap enough that a guy with $20,000 and no morals could build it from his house. That’s the actual AI story the Times should have written.
Being good at deceptive marketing while selling weight-loss and erectile dysfunction drugs online has been a thing since the dawn of email spam. The only novelty here is the tools used to do it. The New York Times just wrapped that up in a neat bow and presented it as the proof of Sam Altman’s big promises for AI.
For what it’s worth, Gallagher has been whining about all this on X, per Futurism’s Dupre:
Though Medvi has yet to respond to our questions, the company’s founder, Gallagher, has spent the last few days on X defending his company. He complained in one post — seemingly in reference to criticism — that “the most low t [testosterone] guys” are “the loudest online” and the “Karens of the internet.” In another post, he wrote that it’s “actually a little crazy the number of people who form a whole opinion from a headline and then publicly wish horrible things will happen.”
Ah yes. The guy complaining about “low t guys” and “karens on the internet” for questioning his “AI business” skills, sure is a trustworthy kind of business person that deserves a NYT puff piece.
The real issue now is what the New York Times plans to do about this. A standard correction noting a few missing details won’t cut it. The entire premise of the article — that this company represents the exciting realization of AI’s business potential — is nonsense. Every element of the narrative is tainted: the growth story is built on deceptive marketing, the product claims are contradicted by the FDA and the manufacturers of the actual drugs, the “$1.8 billion” figure is a projection with no valuation to back it up, and the company is currently facing legal action on multiple fronts. The entire article should be retracted.
The NYT says it “was given access to Medvi’s financials to verify its revenue and profits.” Great. They verified that a company engaged in widespread deceptive practices was, in fact, making money from those deceptive practices. Congrats to the NYT for auditing a snake oil salesman and presenting your findings as if he were an upstanding pharmaceutical salesman.
So to my friends and family members wondering why I haven’t built my own billion-dollar AI company: apparently the missing ingredient wasn’t AI — it was being willing to run a deepfake-powered spam operation selling potentially inert pills to desperate people. The AI just made the lying faster. And the New York Times made one guy appear respectable.
Imagine a newspaper publisher announcing it will no longer allow libraries to keep copies of its paper.
That’s effectively what’s begun happening online in the last few months. The Internet Archive—the world’s largest digital library—has preserved newspapers since it went online in the mid-1990s. The Archive’s mission is to preserve the web and make it accessible to the public. To that end, the organization operates the Wayback Machine, which now contains more than one trillion archived web pages and is used daily by journalists, researchers, and courts.
But in recent months The New York Times began blocking the Archive from crawling its website, using technical measures that go beyond the web’s traditional robots.txt rules. That risks cutting off a record that historians and journalists have relied on for decades. Other newspapers, including The Guardian, seem to be following suit.
For nearly three decades, historians, journalists, and the public have relied on the Internet Archive to preserve news sites as they appeared online. Those archived pages are often the only reliable record of how stories were originally published. In many cases, articles get edited, changed, or removed—sometimes openly, sometimes not. The Internet Archive often becomes the only source for seeing those changes. When major publishers block the Archive’s crawlers, that historical record starts to disappear.
The Times says the move is driven by concerns about AI companies scraping news content. Publishers seek control over how their work is used, and several—including the Times—are now suing AI companies over whether training models on copyrighted material violates the law. There’s a strong case that such training is fair use.
Whatever the outcome of those lawsuits, blocking nonprofit archivists is the wrong response. Organizations like the Internet Archive are not building commercial AI systems. They are preserving a record of our history. Turning off that preservation in an effort to control AI access could essentially torch decades of historical documentation over a fight that libraries like the Archive didn’t start, and didn’t ask for.
If publishers shut the Archive out, they aren’t just limiting bots. They’re erasing the historical record.
Archiving and Search Are Legal
Making material searchable is a well-established fair use. Courts have long recognized it’s often impossible to build a searchable index without making copies of the underlying material. That’s why when Google copied entire books in order to make a searchable database, courts rightly recognized it as a clear fair use. The copying served a transformative purpose: enabling discovery, research, and new insights about creative works.
The Internet Archive operates on the same principle. Just as physical libraries preserve newspapers for future readers, the Archive preserves the web’s historical record. Researchers and journalists rely on it every day. According to Archive staff, Wikipedia alone links to more than 2.6 million news articles preserved at the Archive, spanning 249 languages. And that’s only one example. Countless bloggers, researchers, and reporters depend on the Archive as a stable, authoritative record of what was published online.
The same legal principles that protect search engines must also protect archives and libraries. Even if courts place limits on AI training, the law protecting search and web archiving is already well established.
The Internet Archive has preserved the web’s historical record for nearly thirty years. If major publishers begin blocking that mission, future researchers may find that huge portions of that historical record have simply vanished. There are real disputes over AI training that must be resolved in courts. But sacrificing the public record to fight those battles would be a profound, and possibly irreversible, mistake.
This was extremely wild shit to be happening anywhere, much less in the land of the First Amendment. No sooner had Donald Trump decided it was time to rename the Department of Defense to the Department of War than the head of DoD operations decided it would be sorting news agencies by level of subservience.
Pretending this was all about national security, the Defense Department basically kicked everyone out of the Pentagon’s press office and stated that only those that chose to play by the new rules would be allowed back inside.
Booted: NBC News, the New York Times, NPR. Welcomed back into the fold: OAN, Newsmax, Breitbart. The Pentagon wanted a state-run press, but without having to do all the heavy lifting that comes with instituting a state-run press in the Land of the Free.
Somewhat surprisingly, some of those explicitly invited to partake of the new Defense Department media wing refused to participate. Fox and Newsmax decided to stay out, rather than promise they’d never publish leaked documents. Those choosing to bend the knee were those who never needed this sort of coercion in the first place: One America News (OAN), The Federalist, and far-right weirdos, the Epoch Times. In other words, MAGA-heavy breathers that have never been known for their independence, much less their journalism.
That didn’t stop Hegseth and the department he’s mismanaging from attempting to take a victory lap. And it certainly didn’t stop news agencies like the New York Times from suing over this blatant violation of the First Amendment.
It’s so obvious it only took the NYT four months to secure a win in a federal court (DC) that is positively swamped with litigation generated by Trump’s swamp. (h/t Adam Klasfield)
The decision [PDF] makes it clear in the opening paragraph how this is going to go for the administration and its extremely selective “respect” of enshrined rights and freedoms.
A primary purpose of the First Amendment is to enable the press to publish what it will and the public to read what it chooses, free of any official proscription. Those who drafted the First Amendment believed that the nation’s security requires a free press and an informed people and that such security is endangered by governmental suppression of political speech. That principle has preserved the nation’s security for almost 250 years. It must not be abandoned now.
Amen.
The court notes that in the past, there has been some friction between national security concerns and reporting by journalists. In some cases, the friction has been little more than the government chafing a bit when something has been published that it would rather have kept a secret. In other cases, leaks involving sensitive information have provoked reform efforts on both sides of the equation, seeking to balance these concerns with serving the public interest.
Up until now, any efforts to expel reporters have been limited to backroom bitching. What’s happening now, however, is unprecedented.
Historically, though, even when Department leaders disliked a journalist’s reporting, they did not consider suspending, revoking, or not renewing the journalist’s press credentials in response to that reporting. Julian Barnes, Pete Williams, and Robert Burns—reporters who have spent decades covering the Pentagon—as well as former Pentagon officials, are not aware of the Department ever suspending, revoking, or not renewing a journalist’s credentials due to concern over the safety or security of Department personnel or property or based on the content of their reporting.
This may be new, but the court isn’t willing to make it the “new normal.” It’s the decades of precedent that truly matter, not the vindictive whims of the overgrown toddlers currently holding office.
The Pentagon claims that demanding journalists agree not to “solicit,” much less print data or information not explicitly approved for release by the Defense Department doesn’t reach any further than existing laws governing the handling of classified documents. The court disagrees, noting that the new policy allows the government to conflate the illegal solicitation of classified material with the sort of soliciting — i.e., requests for information, etc. — journalists do every day in hopes of securing something newsworthy.
On top of allowing the government to punish people for things that weren’t previously considered unlawful, the demand for obeisance wasn’t created in a vacuum. Instead, it flowed directly from this entire administration’s constant attacks on the press by the president and pretty much every one in his Cabinet.
The plaintiffs are correct: “The record is replete with undisputed evidence that the Policy is viewpoint discriminatory.” That evidence tells the story of a Department whose leadership has been and continues to be openly hostile to the “mainstream media” whose reporting it views as unfavorable, but receptive to outlets that have expressed “support for the Trump administration in the past.”
The story begins prior to the adoption of the Policy, when—following extensive reporting on Secretary Hegseth’s background and qualifications during his confirmation process—Secretary Hegseth and Department officials “openly complained about reporting they perceive[d] as unfavorable to them and the Department.” Then, in the weeks and months leading up to the issuance of the Policy, Department officials repeatedly condemned certain news organizations—including The Times—for their coverage of the Department. For example, in response to reporting by The Times on Secretary Hegseth’s alleged misuse of the messaging platform Signal, Mr. Parnell posted on X to call out The Times “and all other Fake News that repeat their garbage.” Mr. Parnell decried these news organizations as “Trump-hating media” who “continue[] to be obsessed with destroying anyone committed to President Trump’s agenda.” In other social media posts leading up to the issuance of the Policy, Department officials referred to journalists from The Washington Post as “scum” and called for their “severe punishment” in response to reporting on Secretary Hegseth’s security detail.
It was never about keeping loose lips from sinking ships. It was always about cutting off access to news agencies the administration didn’t like. And once you’ve gotten rid of the critics, you’re left with the functional equivalent of a state-run media, but without the nastiness of having to disappear people into concentration camps or usher them out of their cubicles at gunpoint.
The court won’t let this stand. The new policy violates both the First Amendment and Fifth Amendment (due to the vagueness of its ban on “soliciting” sensitive information). That’s never been acceptable before in this nation. Just because there’s an aspiring tyrant leaning heavily on the Resolute Desk these days doesn’t make it any more permissible.
The Court recognizes that national security must be protected, the security of our troops must be protected, and war plans must be protected. But especially in light of the country’s recent incursion into Venezuela and its ongoing war with Iran, it is more important than ever that the public have access to information from a variety of perspectives about what its government is doing—so that the public can support government policies, if it wants to support them; protest, if it wants to protest; and decide based on full, complete, and open information who they are going to vote for in the next election. As Justice Brandeis correctly observed, “sunlight is the most powerful of all disinfectants.”
The administration will definitely appeal this decision. And it almost definitely will try to bypass the DC Appeals Court and go straight to the Supreme Court by claiming not being able to expel reporters it doesn’t like is some sort of national emergency. It will probably even claim that the fight it picked in Iran justifies the actions it took months before it decided to involve us in the nation’s latest Afghanistan/Vietnam.
But it definitely shouldn’t win. This isn’t some obscure permutation of First Amendment law. This is the government crafting a policy that allows it to decide what gets to be printed and who gets to print it. That’s never been acceptable here. And it never should be.
Last fall, I wrote about how the fear of AI was leading us to wall off the open internet in ways that would hurt everyone. At the time, I was worried about how companies were conflating legitimate concerns about bulk AI training with basic web accessibility. Not surprisingly, the situation has gotten worse. Now major news publishers are actively blocking the Internet Archive—one of the most important cultural preservation projects on the internet—because they’re worried AI companies might use it as a sneaky “backdoor” to access their content.
This is a mistake we’re going to regret for generations.
Nieman Lab reports that The Guardian, The New York Times, and others are now limiting what the Internet Archive can crawl and preserve:
When The Guardian took a look at who was trying to extract its content, access logs revealed that the Internet Archive was a frequent crawler, said Robert Hahn, head of business affairs and licensing. The publisher decided to limit the Internet Archive’s access to published articles, minimizing the chance that AI companies might scrape its content via the nonprofit’s repository of over one trillion webpage snapshots.
Specifically, Hahn said The Guardian has taken steps to exclude itself from the Internet Archive’s APIs and filter out its article pages from the Wayback Machine’s URLs interface. The Guardian’s regional homepages, topic pages, and other landing pages will continue to appear in the Wayback Machine.
The Times has gone even further:
The New York Times confirmed to Nieman Lab that it’s actively “hard blocking” the Internet Archive’s crawlers. At theend of 2025, the Times also added one of those crawlers —archive.org_bot — to itsrobots.txt file, disallowing access to its content.
“We believe in the value of The New York Times’s human-led journalism and always want to ensure that our IP is being accessed and used lawfully,” said a Times spokesperson. “We are blocking the Internet Archive’s bot from accessing the Times because the Wayback Machine provides unfettered access to Times content — including by AI companies — without authorization.”
I understand the concern here. I really do. News publishers are struggling, and watching AI companies hoover up their content to train models that might then, in some ways, compete with them for readers is genuinely frustrating. I run a publication myself, remember.
But blocking the Internet Archive isn’t going to stop AI training. What it will do is ensure that significant chunks of our journalistic record and historical cultural context simply… disappear.
And that’s bad.
The Internet Archive is the most famous nonprofit digital library, and has been operating for nearly three decades. It isn’t some fly-by-night operation looking to profit off publisher content. It’s trying to preserve the historical record of the internet—which is way more fragile than most people comprehend. When websites disappear—and they disappear constantly—the Wayback Machine is often the only place that content still exists. Researchers, historians, journalists, and ordinary citizens rely on it to understand what actually happened, what was actually said, what the world actually looked like at a given moment.
In a digital era when few things end up printed on paper, the Internet Archive’s efforts to permanently preserve our digital culture are essential infrastructure for anyone who cares about historical memory.
And now we’re telling them they can’t preserve the work of our most trusted publications.
Think about what this could mean in practice. Future historians trying to understand 2025 will have access to archived versions of random blogs, sketchy content farms, and conspiracy sites—but not The New York Times. Not The Guardian. Not the publications that we consider the most reliable record of what’s happening in the world. We’re creating a historical record that’s systematically biased against quality journalism.
Yes, I’m sure some will argue that the NY Times and The Guardian will never go away. Tell that to the readers of the Rocky Mountain News, which published for 150 years before shutting down in 2009, or to the 2,100+ newspapers that have closed since 2004. Institutions—even big, prominent, established ones—don’t necessarily last.
As one computer scientist quoted in the Nieman piece put it:
“Common Crawl and Internet Archive are widely considered to be the ‘good guys’ and are used by ‘the bad guys’ like OpenAI,” said Michael Nelson, a computer scientist and professor at Old Dominion University. “In everyone’s aversion to not be controlled by LLMs, I think the good guys are collateral damage.”
That’s exactly right. In our rush to punish AI companies, we’re destroying public goods that serve everyone.
The most frustrating bit of all of this: The Guardian admits they haven’t actually documented AI companies scraping their content through the Wayback Machine. This is purely precautionary and theoretical. They’re breaking historical preservation based on a hypothetical threat:
The Guardian hasn’t documented specific instances of its webpages being scraped by AI companies via the Wayback Machine. Instead, it’s taking these measures proactively and is working directly with the Internet Archive to implement the changes.
And, of course, as one of the “good guys” of the internet, the Internet Archive is willing to do exactly what these publishers want. They’ve always been good about removing content or not scraping content that people don’t want in the archive. Sometimes to a fault. But you can never (legitimately) accuse them of malicious archiving (even if music labels and book publishers have).
Either way, we’re sacrificing the historical record not because of proven harm, but because publishers are worried about what might happen. That’s a hell of a tradeoff.
This isn’t even new, of course. Last year, Reddit announced it would block the Internet Archive from archiving its forums—decades of human conversation and cultural history—because Reddit wanted to monetize that content through AI licensing deals. The reasoning was the same: can’t let the Wayback Machine become a backdoor for AI companies to access content Reddit is now selling. But once you start going down that path, it leads to bad places.
The Nieman piece notes that, in the case of USA Today/Gannett, it appears that there was a company-wide decision to tell the Internet Archive to get lost:
In total, 241 news sites from nine countries explicitly disallow at least one out of the four Internet Archive crawling bots.
Most of those sites (87%) are owned by USA Today Co., the largest newspaper conglomerate in the United States formerly known as Gannett. (Gannett sites only make up 18% of Welsh’s original publishers list.) Each Gannett-owned outlet in our dataset disallows the same two bots: “archive.org_bot” and “ia_archiver-web.archive.org”. These bots were added to the robots.txt files of Gannett-owned publications in 2025.
Some Gannett sites have also taken stronger measures to guard their contents from Internet Archive crawlers.URL searches for the Des Moines Register in the Wayback Machinereturn a message that says, “Sorry. This URL has been excluded from the Wayback Machine.”
A Gannett spokesperson told NiemanLab that it was about “safeguarding our intellectual property” but that’s nonsense. The whole point of libraries and archives is to preserve such content, and they’ve always preserved materials that were protected by copyright law. The claim that they have to be blocked to safeguard such content is both technologically and historically illiterate.
And here’s the extra irony: blocking these crawlers may not even serve publishers’ long-term interests. As I noted in my earlier piece, as more search becomes AI-mediated (whether you like it or not), being absent from training datasets increasingly means being absent from results. It’s a bit crazy to think about how much effort publishers put into “search engine optimization” over the years, only to now block the crawlers that feed the systems a growing number of people are using for search. Publishers blocking archival crawlers aren’t just sacrificing the historical record—they may be making themselves invisible in the systems that increasingly determine how people discover content in the first place.
The Internet Archive’s founder, Brewster Kahle, has been trying to sound the alarm:
“If publishers limit libraries, like the Internet Archive, then the public will have less access to the historical record.”
But that warning doesn’t seem to be getting through. The panic about AI has become so intense that people are willing to sacrifice core internet infrastructure to address it.
What makes this particularly frustrating is that the internet’s openness was never supposed to have asterisks. The fundamental promise wasn’t “publish something and it’s accessible to all, except for technologies we decide we don’t like.” It was just… open. You put something on the public web, people can access it. That simplicity is what made the web transformative.
Now we’re carving out exceptions based on who might access content and what they might do with it. And once you start making those exceptions, where do they end? If the Internet Archive can be blocked because AI companies might use it, what about research databases? What about accessibility tools that help visually impaired users? What about the next technology we haven’t invented yet?
This is a real concern. People say “oh well, blocking machines is different from blocking humans,” but that’s exactly why I mention assistive tech for the visually impaired. Machines accessing content are frequently tools that help humans—including me. I use an AI tool to help fact check my articles, and part of that process involves feeding it the source links. But increasingly, the tool tells me it can’t access those articles to verify whether my coverage accurately reflects them.
I don’t have a clean answer here. Publishers genuinely need to find sustainable business models, and watching their work get ingested by AI systems without compensation is a legitimate grievance—especially when you see how much traffic some of these (usually less scrupulous) crawlers dump on sites. But the solution can’t be to break the historical record of the internet. It can’t be to ensure that our most trusted sources of information are the ones that disappear from archives while the least trustworthy ones remain.
We need to find ways to address AI training concerns that don’t require us to abandon the principle of an open, preservable web. Because right now, we’re building a future where historians, researchers, and citizens can’t access the journalism that documented our era. And that’s not a tradeoff any of us should be comfortable with.
Remember last summer when everyone was freaking out about the explosion of AI-generated child sexual abuse material? The New York Times ran a piece in July with the headline “A.I.-Generated Images of Child Sexual Abuse Are Flooding the Internet.” NCMEC put out a blog post calling the numbers an “alarming increase” and a “wake-up call.” The numbers were genuinely shocking: NCMEC reported receiving 485,000 AI-related CSAM reports in the first half of 2025, compared to just 67,000 for all of 2024.
That’s a big increase! And it would obviously be super concerning if any AI company were finding and detecting so much AI-generated CSAM, especially as we keep hearing that the big AI models (perhaps with the exception of Grok…) have been putting in place safeguards against CSAM generation.
The source of most of those reports? Amazon, which had submitted a staggering 380,000 of them, even though most people don’t tend to think of Amazon as much of an AI company. But, still, it became a six alarm fire about how much AI-generated CSAM Amazon had discovered. There were news stories about it, politicians demanding action, and the general sentiment was that this proved how big the problem was.
Except… it turns out that wasn’t actually what was happening. At all.
Bloomberg just published a deep dive into what was actually going on with Amazon’s reports, and the truth is very, very different from what everyone assumed. According to Bloomberg:
Amazon.com Inc. reported hundreds of thousands of pieces of content last year that it believed included child sexual abuse, whichit found in data gathered to improve its artificial intelligence models. Though Amazon removed the content before training its models, child safety officials said the company has not provided information about its source, potentially hindering law enforcement from finding perpetrators and protecting victims.
Here’s the kicker—and I cannot stress this enough—none of Amazon’s reports involved AI-generated CSAM.
None of its reports submitted to NCMEC were of AI-generated material, the spokesperson added. Instead, the content was flagged by an automatic detection tool that compared it against a database of known child abuse material involving real victims, a process called “hashing.” Approximately 99.97% of the reports resulted from scanning “non-proprietary training data,” the spokesperson said.
What Amazon was actually reporting was known CSAM—images of real victims that already existed in databases—that their scanning tools detected in datasets being considered for AI training. They found it using traditional hash-matching detection tools, flagged it, and removed it before using the data. Which is… actually what you’d want a company to do?
But because it was found in the context of AI development, and because NCMEC’s reporting form has exactly one checkbox that says “Generative AI” with no way to distinguish between “we found known CSAM in our training data pipeline” and “our AI model generated new CSAM,” Amazon checked the box.
And thus, a massive misunderstanding was born.
Again, let’s be clear and separate out a few things here: the fact that Amazon found CSAM (known or not) in its training data is bad. It is a troubling sign of how much CSAM is found in the various troves of data AI companies use for training. And maybe the focus should be on that. Also, the fact that they then reported it to NCMEC and removed it from their training data after discovering it with hash matching is… good. That’s how things are supposed to work.
But the fact that the media (with NCMEC’s help) turned this into “OMG AI generated CSAM is growing at a massive rate” is likely extremely misleading.
For half a year, “Massive Spike In AI-Generated CSAM” is the framing I’ve seen whenever news reports mention those H1 2025 numbers. Even thepress releasefor a Senate bill about safeguarding AI models from being tainted with CSAM stated, “According to the National Center for Missing & Exploited Children, AI-generated material has proliferated at an alarming rate in the past year,” citing the NYT article.
Now we find out from Bloomberg that zero of Amazon’s reports involved AI-generated material; all 380,000 were hash hits to known CSAM. And we have Fallon [McNulty, executive director of the CyberTipline] confirming to Bloomberg that “with the exception of Amazon, the AI-related reports [NCMEC] received last yearcame in ‘really, really small volumes.'”
That is an absolutely mindboggling misunderstanding for everyone — the general public, lawmakers, researchers like me, etc. — to labor under for so long. If Bloomberg hadn’t dug into Amazon’s numbers, it’s not clear to me when, if ever, that misimpression would have been corrected.
She’s not wrong. Nearly 80% of all “Generative AI” CyberTipline reports to NCMEC in the first half of 2025 involved no AI-generated CSAM at all. The actual volume of AI-generated CSAM being reported? Apparently “really, really small.”
Now, to be (slightly?) fair to the NYT, they did run a minor correction a day after their original story noting that the 485,000 reports “comprised both A.I.-generated material and A.I. attempts to create material, not A.I.-generated material alone.” But that correction still doesn’t capture what actually happened. It wasn’t “AI-generated material and attempts”—it was overwhelmingly “known CSAM detected during AI training data vetting.” Those are very different things.
And it gets worse. Bloomberg reports that Amazon’s scanning threshold was set so low that many of those reports may not have even been actual CSAM:
Amazon believes it over-reported these cases to NCMEC to avoid accidentally missing something. “We intentionally use an over-inclusive threshold for scanning, which yields a high percentage of false positives,” the spokesperson added.
So we’ve got reports that aren’t AI-generated CSAM, many of which may not even be CSAM at all. Very helpful.
The frustrating thing is that this kind of confusion wasn’t just entirely predictable—it was predicted! When Pfefferkorn and her colleagues at Stanford published their report about NCMEC’s CSAM reporting system they literally called out the potential confusion in the options of what to check and how platforms would likely over-report stuff in an abundance of caution, because the penalty (both criminally and in reputation) for missing anything is so dire.
Indeed, the form for submitting to the CyberTipline has one checkbox for “Generative AI” that, as Pfefferkorn notes in her letter, can mean wildly different things depending on who’s checking it:
When the meaning of checking a single checkbox is so ambiguous that absent additional information, reports of known CSAM found in AI training data are facially indistinguishable from reports of new AI-generated material (or of text-only prompts seeking CSAM, or of attempts to upload known CSAM as part of a prompt, etc.), and that ambiguity leads to a months-long massive public misunderstanding about the scale of the AI-CSAM problem, then it is clear thatthe CyberTipline reporting form itself needs to change— not just how one particular ESP fills it out.
To its credit NCMEC did respond quickly to Pfefferkorn, and their response is… illuminating. They confirmed they’re working on updating the reporting system, but also noted that Amazon’s reports contained almost no useful information:
all those Amazon reports included minimal data, not even the file in question or the hash value, much less other contextual information about where or how Amazon detected the matching file
As Pfefferkorn put it, Amazon was basically giving NCMEC reports that said “we found something” with nothing else attached. NCMEC says they only learned about the false positives issue last week and are “very frustrated” by it.
Indeed, NCMEC’s boss told Bloomberg:
“There’s nothing then that can be done with those reports,” she said. “Our team has been really clear with [Amazon] that those reports are inactionable.”
There’s plenty of blame to go around here. Amazon clearly should have been more transparent about what they were reporting and why. NCMEC’s reporting form is outdated and creates ambiguity that led to a massive public misunderstanding. And the media (NYT included) ran with alarming numbers without asking obvious questions like “why is Amazon suddenly reporting 25x more than last year and no other AI company is even close?”
But, even worse, policymakers spent six months operating under the assumption that AI-generated CSAM was exploding at an unprecedented rate. Legislation was proposed. Resources were allocated. Public statements were made. All based on numbers that fundamentally misrepresented what was actually happening.
As Pfefferkorn notes:
Nobody benefits from being so egregiously misinformed. It isn’t a basis for sound policymaking (or an accurate assessment of NCMEC’s resource needs) if the true volume of AI-generated CSAM being reported is a mere fraction of what Congress and other regulators believe it is. It isn’t good for Amazon if people mistakenly think the company’s AI products are uniquely prone to generating CSAM compared with other options on the market (such as OpenAI, with its distant-second 75,000 reports during the same time period,per NYT). That impression also disserves users trying to pick safe, responsible AI tools to use; in actuality, per today’s revelations about training data vetting, Amazon is indeed trying to safeguard its models against CSAM. I can certainly think of at least one other AI company that’s been in the news a lot lately that seems to be acting far more carelessly.
None of this means that AI-generated CSAM isn’t a real and serious problem. It absolutely is, and it needs to be addressed. But you can’t effectively address a problem if your data about the scope of that problem is fundamentally wrong. And you especially can’t do it when the “alarming spike” that everyone has been pointing to turns out to be something else entirely.
The silver lining here, as Pfefferkorn points out, is that the actual news is… kind of good? Amazon’s AI models aren’t CSAM-generating machines. The company was actually doing the responsible thing by vetting its training data. And the real volume of AI-generated CSAM reports is apparently much lower than we’ve been led to believe.
But that good news was buried for six months under a misleading narrative that nobody bothered to dig into until Bloomberg did. And that’s a failure of transparency, of reporting systems, and of the kind of basic journalistic skepticism that should have kicked in when one company was suddenly responsible for 78% of all reports in a category.
We’ll see if NCMEC’s promised updates to the reporting form actually address these issues. In the meantime, maybe we can all agree that the next time there’s a 700% increase in reports of anything, it’s worth asking a few questions before writing the “everything is on fire” headline.
Imagine you’re writing an article about a popular policy trend. The trend is expensive to implement, disruptive to normal operations, and—here’s the key part—there’s substantial research showing it doesn’t actually work and can cause other significant problems. How would you structure that article?
One approach: Lead with the evidence. “Despite growing enthusiasm for [policy proposal], studies consistently find it doesn’t accomplish its stated goals.” Put that in paragraph one, maybe paragraph two or three with some lead-up if you’re feeling generous.
Another approach: Spend 13 paragraphs hyping up the trend, listing every conceivable harm it’s meant to address, quoting lawmakers and administrators who support it, and then—only then—casually mention that the evidence shows it doesn’t work.
Mobile phone bans in school and social media bans for kids are increasingly popular around the globe, driven largely by Jonathan Haidt’s bestselling book—which remains a bestseller despite actual experts debunking basically everything in it. So when the paper of record wades into this debate, you’d think they might lead with what the evidence actually shows. You’d think wrong.
The article opens with the traditional moral panic opening, playing up all the fear:
Bullying. Sextortion. Body-shaming. Self-harm. Viralstudent-fight videos. Never-ending newsfeeds. Unhealthy relationships with A.I. chatbots. Teenagers who can’t seem to put down their phones.
Parents and teachers are understandably concerned about social media. For all of the community, creativity and just plain fun kids enjoy online, hazards remain all too frequent, some children’s advocates say.
It’s the greatest-hits compilation of every anxiety adults have projected onto kids and technology for decades (centuries, really). Might as well add “Dungeons & Dragons will make them worship Satan” for completeness.
The piece does eventually ask “can these bans actually help?” But not before spending several more paragraphs cataloging every conceivable harm that’s ever been tangentially associated with social media, strongly implying the tech itself is to blame rather than, you know, humanity. Then it dutifully reports that “lawmakers and schools” see bans as the answer.
Only then—14 paragraphs deep—does the Times get around to mentioning:
Wehave limited researchon whether the bans work. After surveying more than 1,200 students in 30 schools across England, researchers at the University of Birmingham recently reported that cellphone bans did not improve students’ mental well-being.
“Limited research”?
No. We have plenty of research. There’s a comprehensive study in Australia that found no evidence bans helped kids. Multiple reports document actual harms from these bans—including privacy violations and safety issues when kids can’t reach parents during emergencies. It appears that the evidence is just inconvenient for the narrative.
But the Times isn’t done. The article includes a section on how bans “may have drawbacks”—and somehow the main drawback they identify is that bans don’t stop social media companies from doing bad things. Not that the bans don’t work. Not that they create new problems. Just that they don’t magically fix the platforms themselves:
Blanket tech bans can be crude instruments. They may make it harder for many young people to have social media accounts. But they often don’t change the underlying app features that many parents are worried about.
Many popular apps use powerful attention-hacking techniques that can hook young people, said Julia Powles, an Australian researcher who is the executive director of the U.C.L.A. Institute for Technology, Law and Policy. This keeps users online longer, she notes, and makes the companies more money from advertising.
This completely misses the point—which, as danah boyd has repeatedly explained, is that adults are confusing risks with harms. Many things are risky. Some can lead to harm. But we generally deal with risky things by teaching people how to manage those risks.
The response to potential harms from social media shouldn’t be to demand bans. It should be teaching kids how to navigate these spaces appropriately—how to recognize manipulation, how to minimize risks, what to do when something goes wrong. Instead, we hide it. We ban it. We shove it under the rug and pretend that if we just keep this scary thing away from kids, they’ll somehow be fine once the ban lifts.
And thus, we get the worst of everything. For every ban out there, kids will find their ways around them. Often, that will involve doing things surreptitiously, in places with fewer controls and less ability for parents and teachers to properly instruct kids how to use those tools appropriately. It actually puts kids in more danger by pretending that if we just “ban” places for them to communicate, that they’ll just become perfect little kids who never look elsewhere.
The Times had a chance here to actually inform the debate—to lead with what the evidence shows, to explain the tradeoffs, to challenge the reflexive push for bans. Instead, they wrote 13 paragraphs of pure moral panic before mentioning that these policies don’t work, then immediately pivoted back to fearmongering about “attention-hacking techniques.”
This all just feeds the moral panic. It gives politicians and administrators cover to implement bans that won’t help kids but will absolutely create new problems. And when those bans inevitably fail, the Times will probably write another breathless piece wondering why kids are still struggling—while once again burying the fact that we never actually tried teaching them how to navigate these spaces in the first place.
A federal magistrate judge just ordered that the private ChatGPT conversations of 20 million users be handed over to the lawyers for dozens of plaintiffs, including news organizations. Those 20 million people weren’t asked. They weren’t notified. They have no say in the matter.
Last week, Magistrate Judge Ona Wang ordered OpenAI to turn over a sample of 20 million chat logs as part of the sprawling multidistrict litigation where publishers are suing AI companies—a mess of consolidated cases that kicked off with the NY Times’ lawsuit against OpenAI. Judge Wang dismissed OpenAI’s privacy concerns, apparently convinced that “anonymization” solves everything.
Even if you hate OpenAI and everything it stands for, and hope that the news orgs bring it to its knees, this should scare you. A lot. OpenAI had pointed out to the judge a week earlier that this demands from the news orgs would represent a massive privacy violation for ChatGPT’s users.
News Plaintiffs demand that OpenAI hand over the entire 20M log sample “in readily searchable format” via a “hard drive or [] dedicated private cloud.” ECF 656 at 3. That would include logs that are neither relevant nor responsive—indeed, News Plaintiffs concede that at least 99.99% of the logs are irrelevant to their claims. OpenAI has never agreed to such a process, which is wildly disproportionate to the needs of the case and exposes private user chats for no reasonable litigation purpose. In a display of striking hypocrisy, News Plaintiffs disregard those users’ privacy interests while claiming that their own chat logs are immune from production because “it is possible” that their employees “entered sensitive information into their prompts.” ECF 475 at 4. Unlike News Plaintiffs, OpenAI’s users have no stake in this case and no opportunity to defend their information from disclosure. It makes no sense to order OpenAI to hand over millions of irrelevant and private conversation logs belonging to those absent third parties while allowing News Plaintiffs to shield their own logs from disclosure.
OpenAI offered a much more privacy-protective alternative: hand over only a targeted set of logs actually relevant to the case, rather than dumping 20 million records wholesale. The news orgs fought back, but their reply brief is sealed—so we don’t get to see their argument. The judge bought it anyway, dismissing the privacy concerns on the theory that OpenAI can simply “anonymize” the chat logs:
Whether or not the parties had reached agreement to produce the 20 million Consumer ChatGPT Logs in whole—which the parties vehemently dispute—such production here is appropriate. OpenAI has failed to explain how its consumers’ privacy rights are not adequately protected by: (1) the existing protective order in this multidistrict litigation or (2) OpenAI’s exhaustive de-identification of all of the 20 million Consumer ChatGPT Logs.
The judge then quotes the news orgs’ filing, noting that OpenAI has already put in this effort to “deidentify” the chat logs.
Both of those supposed protections—the protective order and “exhaustive de-identification”—are nonsense. Let’s start with the anonymization problem, because it shows a stunning lack of understanding about what it means to anonymize data sets, especially AI chatlogs.
We’ve spent years warning people that “anonymized data” is a gibberish term, used by companies to pretend large collections of data can be kept private, when that’s just not true. Almost any large dataset of “anonymized” data can have significant portions of the data connected back to individuals with just a little work. Researchers re-identified individuals from “anonymized” AOL search queries, from NYC taxi records, from Netflix viewing histories—the list goes on. Every time someone shows up with an “anonymized” dataset, researchers show ways to re-identify people in the dataset.
And that’s even worse when it comes to ChatGPT chat logs, which are likely to be way more revealing that previous data sets where the inability to anonymize data were called out. There have been plenty of reports of just how much people “overshare” with ChatGPT, often including incredibly private information.
Back in August, researchers got their hands on just 1,000 leaked ChatGPT conversations and talked about how much sensitive information they were able to glean from just that small number of chats.
Researchers downloaded and analyzed 1,000 of theleaked conversations,spanning over 43 million words. Among them, they discovered multiple chats that explicitly mentioned personally identifiable information (PII), such as full names, addresses, and ID numbers.
With that level of PII and sensitive information, connecting chats back to individuals is likely way easier than in previous cases of connecting “anonymized” data back to individuals.
And that was with just 1,000 records.
Then, yesterday as I was writing this, the Washington Post revealed that they had combed through 47,000 ChatGPT chat logs, many of which were “accidentally” revealed via ChatGPT’s “share” feature. Many of them reveal deeply personal and intimate information.
Users often shared highly personal information with ChatGPT in the conversations analyzed by The Post, including details generally not typed into conventional search engines.
People sent ChatGPT more than 550 unique email addresses and 76 phone numbers in the conversations. Some are public, but others appear to be private, like those one user shared for administrators at a religious school in Minnesota.
Users asking the chatbot to draft letters or lawsuits on workplace or family disputes sent the chatbot detailed private information about the incidents.
There are examples where, even if the user’s official details are redacted, it would be trivial to figure out who was actually doing the chats:
If you can’t see that, it’s a chat with ChatGPT, redacted by the Washington post saying:
User my name is [name redacted] my husband name [name redacted] is threatning me to kill and not taking my responsibities and trying to go abroad […] he is not caring us and he is going to kuwait and he will give me divorce from abroad please i want to complaint to higher authgorities and immigrition office to stop him to go abroad and i want justice please help
ChatGPT Below is a formal draft complaint you can submit to the Deputy Commissioner of Police in [redacted] addressing your concerns and seeking immediate action:
That seems like even if you “anonymized” the chat by taking off the user account details, it wouldn’t take long to figure out whose chat it was, revealing some pretty personal info, including the names of their children (according to the Post).
And WaPo reporters found that by starting with 93,000 chats, then using tools do an analysis of the 47,000 in English, followed by human review of just 500 chats in a “random sample.”
Now imagine 20 million records. With many, many times more data, the ability to cross-reference information across chats, identify patterns, and connect seemingly disconnected pieces of information becomes exponentially easier. This isn’t just “more of the same”—it’s a qualitatively different threat level.
Even worse, the judge’s order contains a fundamental contradiction: she demands that OpenAI share these chatlogs “in whole” while simultaneously insisting they undergo “exhaustive de-identification.” Those two requirements are incompatible.
Real de-identification would require stripping far more than just usernames and account info—it would mean redacting or altering the actual content of the chats, because that content is often what makes re-identification possible. But if you’re redacting content to protect privacy, you’re no longer handing over the logs “in whole.” You can’t have both. The judge doesn’t grapple with this contradiction at all.
Yes, as the judge notes, this data is kept under the protective order in the case, meaning that it shouldn’t be disclosed. But protective orders are only as strong as the people bound by them, and there’s a huge risk here.
Looking at the docket, there are a ton of lawyers who will have access to these files. The docket list of parties and lawyers is 45 pages long if you try to print it out. While there are plenty of repeats in there, there have to be at least 100 lawyers and possibly a lot more (I’m not going to count them, and while I asked three different AI tools to count them, each gave me a different answer).
That’s a lot of people—many representing entities directly hostile to OpenAI—who all need to keep 20 million private conversations secret.
That’s not even getting into the fact that handling 20 million chat logs is a difficult task to do well. I am quite sure that among all the plaintiffs and all the lawyers, even with the very best of intentions, there’s still a decent chance that some of the content could leak (and it could, in theory, leak to some of the media properties who are plaintiffs in the case).
And, as OpenAI properly points out, its users whose data is at risk here have no say in any of this. They likely have no idea that a ton of people may be about to get an intimate look at what they thought were their private ChatGPT chats.
OpenAI is unaware of any court ordering wholesale production of personal information at this scale. This sets a dangerous precedent: it suggests that anyone who files a lawsuit against an AI company can demand production of tens of millions of conversations without first narrowing for relevance. This is not how discovery works in other cases: courts do not allow plaintiffs suing Google to dig through the private emails of tens of millions of Gmail users irrespective of their relevance. And it is not how discovery should work for generative AI tools either.
The judge had cited a ruling in one of Anthropic’s cases, but hadn’t given OpenAI a chance to explain why the ruling in that case didn’t apply here (in that one, Anthropic had agreed to hand over the logs as part of negotiations with the plaintiffs, and OpenAI gets in a little dig at its competitor, pointing out that it appears Anthropic made no effort to protect the privacy of its users in that case).
There have, as Daphne Keller regularly points out, always been challenges between user privacy and platform transparency. But this goes well beyond that familiar tension. We’re not talking about “platform transparency” in the traditional sense—publishing aggregated statistics or clarifying moderation policies. This is 20 million complete chatlogs, handed over “in whole” to dozens of adversarial parties and their lawyers. The potential damage to the privacy rights of those users could be massive.
Earlier today we wrote about Trump’s extraordinary admission that he was basing military deployment decisions on old Fox News footage and lies from his advisors. But there’s an even more damning story here: how that revelation almost never saw the light of day because of journalistic cowardice.
The smoking gun quote came from Trump’s phone interview with NBC’s Yamiche Alcindor:
“I spoke to the governor, she was very nice,” Trump said. “But I said, ‘Well wait a minute, am I watching things on television that are different from what’s happening? My people tell me different.’ They are literally attacking and there are fires all over the place…it looks like terrible.”
This is an absolutely nuclear quote.
But note that we linked to the local KGW affiliate report on it and not NBC’s.
And that’s because NBC didn’t even mention the quote at all in its own coverage. As Dan Froomkin highlighted in his article about all this, NBC ran two stories by Alcindor (with Alexandra Marquez) about her interview with Trump, neither of which mentioned that bombshell of a quote.
Instead, it was only because NBC apparently sent the full transcript to affiliates that Evan Watson at KGW picked it up and ran a story about it.
But that raises a ton of questions, including how could NBC and Alcindor not see this as a story? And what is wrong with the mainstream media that it basically skipped over this?
The quote is devastating. It reveals a president who is either completely detached from reality, easily manipulated by advisors feeding him false information, or being deliberately deceived by old Fox News footage (as we now know was happening). It raises fundamental questions about who is actually running the country and whether the person with access to nuclear codes can distinguish between television clips from five years ago and reality. As we detailed yesterday, this quote reveals everything about how Trump ended up threatening military action against an American city based on five-year-old Fox News b-roll.
NBC’s failure to see the story in this is journalistic malpractice of the highest order. When the President admits he can’t tell the difference between Fox News b-roll and reality, that’s not a throwaway line—it’s the story.
But it’s also part of a much larger pattern of media cowardice that’s actively damaging public trust in journalism. The problem isn’t just burying important quotes—it’s the widespread adoption of “view from nowhere” reporting that treats even the most basic facts as matters of debate.
Take this astounding example from a recent New York Times piece about Trump’s use of military force against boats in the Caribbean.
Some legal expertshave called it a crime to summarily kill civilians not directly taking part in hostilities, even if they are believed to be smuggling drugs.
“Some legal experts?” Are you kidding me? Summarily executing civilians is a war crime under international law. This isn’t a matter of debate among competing schools of legal thought. There isn’t another camp of legal experts arguing that, actually, murdering civilians is totally fine. The Times is creating false balance where none exists, making it sound like there’s some reasonable disagreement about whether mass murder constitutes a crime.
Or consider this gem from CNN, fact-checking Trump’s claim that he reduced prescription drug prices by 1500%:
Trump has unveiled a number of moves aimed at cutting drug prices in recent months, but he has yet to move the needle on reducing costs – much less slashing them by 1,500%, which is mathematically impossible,experts say.
Experts say? You need experts to tell you that 1500% is more than 100%? This is elementary school math. A 100% reduction means something is free. A 1500% reduction would mean pharmaceutical companies are paying you a decent sum of money to take their pills. You don’t need to consult the National Academy of Sciences to determine this is bullshit—you need to remember fourth grade.
This kind of reporting is journalistic malpractice disguised as objectivity. When reporters feel compelled to add “experts say” to basic mathematical facts or treat war crimes as matters of legitimate debate, they’re not being neutral—they’re actively misleading their audience into believing basic facts are up for debate among “experts.”
The pattern is clear: mainstream media has become so terrified of appearing biased that they’ve abandoned their basic responsibility to clearly communicate truth to the public. They’d rather hide behind the false comfort of “some say” and “experts disagree” than plainly state obvious facts.
This isn’t objectivity—it’s cowardice. And it’s precisely why trust in media continues to crater.
There’s an old joke in the journalism field (with disputes over where it originated from) but the line is “if one person says it’s raining and another says it’s not, the journalist should look outside and report the truth” rather than suggesting whether or not it’s raining is a matter of dispute.
We’re seeing the opposite from the mainstream media these days.
When the President of the United States admits he can’t distinguish between television and reality, that’s not a “both sides” story, or a cute anecdote not worth mentioning. When someone claims to have reduced costs by 1500%, that’s not a matter requiring expert consultation—it’s a mathematical impossibility. When military officials discuss summarily executing civilians, that’s not a policy debate—it’s war crimes.
The public deserves better than this mealy-mouthed nonsense. They deserve reporters who can recognize when they’re witnessing something extraordinary and have the courage to say so clearly. They deserve news organizations that understand the difference between false balance and actual journalism.
Instead, we get reporters who bury the most important quotes of their own interviews and editors who think basic arithmetic requires expert verification. Is it any wonder people are losing faith in institutions that seem incapable of simply stating reality on its own terms?
The media keeps wondering why trust in journalism is at historic lows. Here’s a thought: maybe it’s because when the President reveals he’s making military decisions based on old Fox News footage and lies from his advisors, the reporter who got that admission decides it’s not worth mentioning. Or maybe it’s because the likes of CNN and the NY Times are so worried about angry people attacking them for calling bullshit on the President that they have to cower behind “experts say” on basic objective facts.
That’s not journalism. That’s stenography. And the American people can tell the difference, even when their media apparently cannot.
In what may be a first in American legal history, a sitting president just had his lawsuit struck down by a federal judge before the defendants even had a chance to respond.
Judge Steven Merryday didn’t wait for a motion to dismiss. He didn’t wait for the defendants to file an answer. Four days after Donald Trump’s lawyers filed their 85-page tantrum masquerading as a defamation complaint against the New York Times and Penguin Random House, Merryday struck it sua sponte—essentially telling the President of the United States that his legal filing was so fundamentally defective it wasn’t worth the court’s time.
Sua sponte dismissals are extraordinarily rare. Judges typically bend over backwards to let even the most questionable complaints proceed to motion practice. The fact that a federal judge took the unusual step of striking a complaint without any prompting from defendants signals just how egregiously improper Trump’s filing was.
Last week, we told you about the ridiculously dopey lawsuit that Donald Trump had filed against Penguin Random House, the NY Times, and some reporters over… something. It wasn’t quite clear. But the lawsuit spent many, many pages fluffing Donald Trump’s ego and suggesting that the mere endorsement in the NY Times of Kamala Harris was election interference and suggested that it would break all the laws to criticize Dear Leader Donald J. Trump.
The complaint also betrayed a fundamental misunderstanding of defamation law’s “actual malice” standard and bore hallmarks that led many observers to suspect it was AI-generated—a theory that gains credibility when you read Judge Merryday’s scathing analysis of its contents.
The venue choice was transparently strategic. Trump forum-shopped his way to the Tampa division of the Middle District of Florida despite having no meaningful connection there—Mar-a-Lago is in the Southern District, and the defendants are based in New York. The complaint’s assertion that venue was proper because defendants “sell newspapers and books” in the district was laughably weak.
The real reason was likely that four of the five regular judges in that division were Trump appointees. But Trump’s luck ran out when the case landed on the docket of Judge Steven Merryday (who is on senior status), a no-nonsense Bush Sr. appointee who clearly wasn’t impressed by the presidential plaintiff.
As every member of the bar of every federal court knows (or is presumed to know), Rule 8(a), Federal Rules of Civil Procedure, requires that a complaint include “a short and plain statement of the claim showing that the pleader is entitled to relief.” Rule 8(e)(1) helpfully adds that “[e]ach averment of a pleading shall be simple, concise, and direct.” Some pleadings are necessarily longer than others. The difference likely depends on the number of parties and claims, the complexity of the governing facts, and the duration and scope of pertinent events. But both a shorter pleading and a longer pleading must comprise “simple, concise, and direct” allegations that offer a “short and plain statement of the claim.” Rule 8 governs every pleading in a federal court, regardless of the amount in controversy, the identity of the parties, the skill or reputation of the counsel, the urgency or importance (real or imagined) of the dispute, or any public interest at issue in the dispute.
In this action, a prominent American citizen (perhaps the most prominent American citizen) alleges defamation by a prominent American newspaper publisher (perhaps the most prominent American newspaper publisher) and by several other corporate and natural persons. Alleging only two simple counts of defamation, the complaint consumes eighty-five pages. Count I appears on page eighty, and Count II appears on page eighty-three. Pages one through seventy-nine, plus part of page eighty, present allegations common to both counts and to all defendants. Each count alleges a claim against each defendant and, apparently, each claim seeks the same remedy against each defendant.
But the judge doesn’t mince words about how “improper” the complaint is beyond just the length:
Even under the most generous and lenient application of Rule 8, the complaint is decidedly improper and impermissible. The pleader initially alleges an electoral victory by President Trump “in historic fashion” — by “trouncing” the opponent — and alludes to “persistent election interference from the legacy media, led most notoriously by the New York Times.” The pleader alludes to “the halcyon days” of the newspaper but complains that the newspaper has become a “full-throated mouthpiece of the Democrat party,” which allegedly resulted in the “deranged endorsement” of President Trump’s principal opponent in the most recent presidential election. The reader of the complaint must labor through allegations, such as “a new journalistic low for the hopelessly compromised and tarnished ‘Gray Lady.’” The reader must endure an allegation of “the desperate need to defame with a partisan spear rather than report with an authentic looking glass” and an allegation that “the false narrative about ‘The Apprentice’ was just the tip of Defendants’ melting iceberg of falsehoods.” Similarly, in one of many, often repetitive, and laudatory (toward President Trump) but superfluous allegations, the pleader states, “‘The Apprentice’ represented the cultural magnitude of President Trump’s singular brilliance, which captured the [Z]eitgeist of our time.”
And also points out how “tedious” the complaint is and points out that a civil complaint is no place for ranting and raving about how mean people are to you, with the main target being the PR value over having a legitimate complaint:
As every lawyer knows (or is presumed to know), a complaint is not a public forum for vituperation and invective — not a protected platform to rage against an adversary. A complaint is not a megaphone for public relations or a podium for a passionate oration at a political rally or the functional equivalent of the Hyde Park Speakers’ Corner.
That’s basically: “your complaint is the legal equivalent of the guy screaming out conspiracy theories on the street corner.”
The judge, as he should, gives Trump 28 days to amend the complaint, which is likely to happen. Whether or not his lawyers can actually follow the local rules and properly state a claim will remain only conjecture until that time.
Meanwhile, Trump appeared wholly unaware that the case was tossed on Friday while meeting with the press. He started bragging about the case and when ABC News reporter Jonathan Karl pointed out that the case had been tossed, Trump responded “I’m winning, I’m winning the cases.” He’s not.
TRUMP: That's why I sued the New York Times two days ago for a lot of moneyKARL: A judge just threw that outTRUMP: I'm winning. I'm winning the cases.
The disconnect between Trump’s perception and legal reality perfectly encapsulates his approach to litigation: file theatrical lawsuits designed more for headlines than legal success, then either attack judges or (as here) just deny reality when courts treat them as actual legal documents that must follow rules. It’s a pattern we’ve seen repeatedly—lawsuits that work better as press releases than as instruments of justice.
Having a President who operates in an alternate reality where judicial smackdowns count as victories is, to put it mildly, concerning. But these days, it’s just a Friday.