Getty Images

Platformer

AI “plagiarism engines” like Perplexity cannot be the future of the web

We can still have the internet we want — but we have to try new business models.

Casey Newton

6/21/24 7:52AM

For a while now, I’ve been gloomy about the state of the web. Plagiarism engines like Perplexity and Arc Search have attracted millions of users by ripping off other people’s work, depriving publishers of the traffic and advertising revenue that once sustained them. The results have been successful enough that Google is following them.

Today, I want to talk about a more positive vision for the future of the internet — one where AI companies and creators work hand in hand to grow the web again, sharing the wealth they create with one another.

Before I get there, though, it’s worth taking a moment to reflect on how bad the status quo has gotten.

Earlier this month, Forbes noticed that Perplexity had been stealing its journalism. The AI startup had taken a scoop about Eric Schmidt’s new drone project and repurposed it for its new “pages” product, which creates automated book-report style web pages based on user prompts. Perplexity had apparently decided to take Forbes’ reporting to show off what its plagiarism can do.

Here’s Randall Lane, Forbes’ chief content officer, in a blog post.

“Not just summarizing (lots of people do that), but with eerily similar wording, some entirely lifted fragments — and even an illustration from one of Forbes’ previous stories on Schmidt,” noted “More egregiously, the post, which looked and read like a piece of journalism, didn’t mention Forbes at all, other than a line at the bottom of every few paragraphs that mentioned “sources,” and a very small icon that looked to be the “F” from the Forbes logo – if you squinted. [...]

Perplexity then sent this knockoff story to its subscribers via a mobile push notification. It created an AI-generated podcast using the same (Forbes) reporting — without any credit to Forbes, and that became a YouTube video that outranks all Forbes content on this topic within Google search.

Any reporter who did what Perplexity did would be drummed out of the journalism business.

Any reporter who did what Perplexity did would be drummed out of the journalism business. But CEO Aravind Srinivas attributed the problem here to “rough edges” on a newly released product, and promised attribution would improve over time. “We agree with the feedback you've shared that it should be a lot easier to find the contributing sources and highlight them more prominently,” he wrote in an X post.

In person, Srivinas can come across as earnest and a bit naive, as I learned when he came on Hard Fork in February. But any notion that Perplexity’s problems stem from a simple misunderstanding was dashed this week when Wired published an investigation into how the company sources answers for users’ queries. In short, Wired found compelling evidence that Perplexity is ignoring the Robots Exclusion Protocol, which publishers and other websites use to grant or deny permissions to automated crawlers and scrapers.

Here are Dhruv Mehrotra and Tim Marchman:

Until earlier this week, Perplexity published in its documentation a link to a list of the IP addresses its crawlers use—an apparent effort to be transparent. However, in some cases, as both Wired and Knight were able to demonstrate, it appears to be accessing and scraping websites from which coders have attempted to block its crawler, called Perplexity Bot, using at least one unpublicized IP address. The company has since removed references to its public IP pool from its documentation. [...]

Wired verified that the IP address in question is almost certainly linked to Perplexity by creating a new website and monitoring its server logs. Immediately after a Wired reporter prompted the Perplexity chatbot to summarize the website's content, the server logged that the IP address visited the site. This same IP address was first observed by Knight during a similar test.

Forbes sent Perplexity a cease-and-desist letter, and I imagine it won’t be the last publisher to do so. There are open legal questions about whether copyrighted material can be used to train large language models or answer chatbot queries, but I see no legal way Perplexity can get away with one of its other core techniques for building pages: using copyrighted images from Getty, the Wall Street Journal, Forbes and others. You simply are not allowed to re-publish other people’s copyrighted photos and illustrations without permission, even if your plagiarism engine is new and has “rough edges.”

Perhaps Perplexity will clean up its act; once it came under fire, the company ran to Semafor to promise that it is “working on” deals with publishers. In the meantime, though, I’ve come to think of it as the Clearview AI of generative artificial intelligence companies: scraping billions of pieces of data without permission and daring courts to stop it.

Like Clearview, Perplexity’s core innovation is ethical rather than technical. In the recent past, it would have been considered bad form to steal and repurpose journalism at scale. Perplexity is making a bet that the advent of generative AI has somehow changed the moral calculus to its benefit.

“I think we need to work together to build all these things, rather than trying to see it as, hey, like you’re taking my stuff and using it,” Srinivas told us in February.

But then he just kept taking everyone’s stuff and using it. The working together part, I guess, is meant to come later.

II.

One path forward for the web, as I shared on a recent episode of Search Engine, is the Fediverse. Decentralized, federated apps; portable identities and follower graphs; permissionless innovation on open protocols: this is a way journalists can once again begin to build audiences — stable ones! — rather than simply courting traffic. This is a years-long project, and I can only barely see the outlines of it taking shape. But it’s an appealing alternative to a world where all content is subsumed into a large language model and accessed by an opaque and proprietary set of algorithms.

But this is a long-term solution, and a partial one. And it carries with it the embedded assumption that today’s AI systems cannot be reshaped in ways that actually grow the web, and pay for the labor of the people who make it. The Fediverse is about giving up on the consumer internet as we know it today — the big walled gardens, the metastasizing LLMs — and trying to build something different.

Tim O’Reilly is thinking differently. As a publisher, investor, and open source advocate, O’Reilly sits at the intersection of many of the business problems and opportunities presented by AI. On Tuesday, he offered his solution to parasitic companies like Perplexity: developing new business models for AI companies that pay creators based on the amount of material that the companies use.

O’Reilly is starting with his own publishing business, sharing a portion of subscription revenue with (or paying a fixed fee to) authors when it uses AI to generate summaries, test questions, translations, or other derivative works based on their writing.

He concludes:

When someone reads a book, watches a video, or attends a live training, the copyright holder gets paid. Why should derivative content generated with the assistance of AI be any different? Accordingly, we have built tools to integrate AI-generated products directly into our payment system. This approach enables us to properly attribute usage, citations, and revenue to content and ensures our continued recognition of the value of our authors’ and teachers’ work.

And if we can do it, we know that others can too.

To O’Reilly, this view of AI is a natural extension of the modern web, which is built on what he calls an “architecture of participation.” The earlier web consisted of giant walled gardens like AOL and MSN, which sought to keep as much activity within their own borders as possible. In this view, companies like Google, OpenAI, and Perplexity are all competing to become the next AOL. It is a vision in which most of the benefits of AI are reaped by a very small number of companies.

“Only the most short-term of business advantage can be found by drying up the river AI companies drink from.”

But this would be a mistake, he writes, if only because the current AI business models are ultimately self-defeating. “If the long-term health of AI requires the ongoing production of carefully written and edited content — as the currency of AI knowledge certainly does — only the most short-term of business advantage can be found by drying up the river AI companies drink from,” O’Reilly writes. “Facts are not copyrightable, but AI model developers standing on the letter of the law will find cold comfort in that if news and other sources of curated content are driven out of business.”

We know that AI companies are running out of data to train their frontier models on. Given that fact, it seems ludicrous that companies like Perplexity are building systems that all but ensure they will have less data to train on in the future.

O’Reilly is taking the opposite approach. And while it remains to be seen whether the average writer on his platform benefits meaningfully from AI royalties, if nothing else he has gotten the incentive structure right. Pay people to create high-quality writing and other content; use that content with permission to train powerful AI systems; and share the wealth that those systems create to fund and incentivize the production of further high-quality writing.

If Srinivas meant it when he said he “we need to work together to build all these things,” he can now look to O’Reilly for a powerful example of what working together actually looks like.

Casey Newton writes Platformer, a daily guide to understanding social networks and their relationships with the world. This piece was originally published on Platformer.

Rani Molla2h

Report: Tesla to build solar factory near Houston

Tesla is planning to build its solar panel manufacturing plant — an endeavor that could add up to $50 billion in value to its energy business — near Houston, Texas, Electrek reports. The plant would be located on the same site as its Megafactory, which builds Megapack battery systems.

The solar plant is part of Tesla and SpaceX’s goal of eventually putting solar-powered data centers in space.

On the company’s fourth-quarter earnings call, CEO Elon Musk said Tesla was “going to work towards getting 100 gigawatts a year of solar cell production, integrating across the entire supply chain from raw materials all the way to finished solar panels.”

At the time, the news had sent shares of First Solar down, but subsequent reports suggest Tesla is unlikely to compete directly with the country’s leading photovoltaic panel maker, instead using much of that production internally.

Exclusive: Tesla (TSLA) is building its giant solar panel factory in Houston

Rani Molla2h

Anthropic hires former OpenAI member and Tesla AI director Andrej Karpathy

Andrej Karpathy — a founding member of OpenAI, Tesla’s director of AI from 2017 to 2022, and the man responsible for the term “vibe coding” — is doing what many in tech are doing right now: heading to greener pastures at Anthropic.

Anthropic, which is slated to go public this year, recently raised money at a $950 billion valuation, making it more valuable than OpenAI and nearly as valuable as Tesla.

Personal update: I've joined Anthropic. I think the next few years at the frontier of LLMs will be especially formative. I am very excited to join the team here and get back to R&D. I remain deeply passionate about education and plan to resume my work on it in time.
— Andrej Karpathy (@karpathy) May 19, 2026

TURN SIGNAL

Confessions of a (former) robotaxi hater

Rani Molla

“governments around the world will not allow Apple junk fees to stand”

Rani Molla5h

Epic Games has returned “Fortnite” to the Apple App Store globally, after the video game maker signaled confidence in its ongoing lawsuit with the iPhone maker. In a press release Tuesday, the company wrote:

“Fortnite is returning to the App Store now because we are confident that once Apple is forced to show its costs, governments around the world will not allow Apple junk fees to stand.

We will continue to challenge Apple’s anticompetitive App Store practices of banning alternative app stores and competition in payments.”

Late last year, an appeals court partly reversed sanctions against Apple but upheld the contempt finding and an injunction forcing Apple to permit outside payment options. “Fortnite” returned to the US App Store a year ago.

The suit began in 2020 over Apple’s mandatory 30% commission on in-app purchases and its refusal to allow third-party payment processors or alternative app stores on its mobile devices.

Rani Molla23h

Meta to lay off 8,000 employees, move 7,000 to new initiatives related to AI

On Wednesday, Reuters reported Meta plans to lay off about 8,000 employees in three batches and move another 7,000 employees to “new initiatives related to AI workflows.” The company also plans to “eliminate managerial roles,” though Reuters did not specify how many.

Reuters had previously reported the number and date of the layoffs, but details of the restructuring come from a new internal document from the company’s head of human resources. The cuts come as Meta tries to balance its enormous capex budget of $125 billion to $145 billion this year, as it builds out its AI infrastructure.

As of the company’s last earnings report, its headcount was 77,986.

Exclusive: Meta lays out plans for May 20 layoffs, restructuring, internal document says