You Already Paid the Subscription. Now They Want Your Data for AI Too

I couldn't take my eyes away from my book for weeks. There I was, every night, reading something jaw-dropping.

The book I was reading was Careless People, where an ambitious lawyer works at Meta for years, seeing the ugly truth behind the tech company.

I always knew there was a price to pay for using free software. We all know free products don't really exist, and it wasn't exactly a shock when years later it turned out our data was being sold to the highest bidder.

Meta never charged anyone and built one of the most valuable companies in the world on the back of what people posted, liked, and clicked. Google didn't charge for Gmail, and for over a decade it scanned the contents of your inbox to target the ads around it, a practice it only stopped in 2017 after years of criticism.

Harvard professor Shoshana Zuboff gave this a name in her 2019 book The Age of Surveillance Capitalism. She described how companies learned to treat human behavior itself as a raw material, something to extract, refine, and sell, whether or not anyone was paying a subscription fee.

"Forget the cliché that if it’s free, 'You are the product.' You are not the product; you are the abandoned carcass. The 'product' derives from the surplus that is ripped from your life."

— Shoshana Zuboff, The Age of Surveillance Capitalism

Both Google and Facebook were honest about the trade, in the sense that it was written into terms of service somewhere. Both also depended on most people never reading it. The deal was simple enough once you saw it: free software runs on data, because there's nothing else for it to run on.

But paid software was supposed to be the escape from that deal.Well, that didn't last long.

Latching onto paid data

You paid the invoice, and your data wouldn't be abused or used for marketing. That was supposed to be the whole pitch of enterprise software for thirty years.

Plenty of paid tools always collected usage analytics: which clicks you made, which features you opened, how long a session lasted. Fine. Most people accepted that as the ordinary cost of using a hosted product.

Training a generative model on your content is a different thing entirely. And making it default on is as bad as when Apple added U2 to my iPhone (it's still haunting me to this day).

It means the actual words your team wrote, and the specific ways your customers described their problems, get absorbed into a system that can reproduce those patterns later.

That can happen for other customers, in other contexts, long after you've forgotten you agreed to anything. It's a bigger thing than "we track how you use the app." It's been sliding into the same "by default" bucket anyway, without most customers noticing the size of the jump.

Increasingly, that line between paid and protected is getting blurry, and the reason? AI.

Six companies, one shape

If you thought this was going to be a "oh it only happened once" kinda story, you'd be wrong, unfortunately.

Zoom was an early example of this term updating scenario.

In March 2023, the company updated its terms to give itself the right to use customer audio, video, and chat content to train its AI models. The change went mostly unnoticed until a Hacker News post surfaced it in August, and the backlash was immediate.

Zoom published a follow-up promising it wouldn't use customer content to train its models without consent.

The consent mechanism it shipped told the real story: a pre-checked box for meeting admins that had to be actively unchecked, and for regular participants, a pop-up that offered a choice between accepting or leaving the meeting.

Zoom's opt-out method. Source: Zoom

Legal analysts pointed out that neither one looks much like consent under European privacy law, and Zoom kept adjusting the terms for months afterward.

Slack followed a similar arc.

Its privacy principles, in place since at least September 2023, said the company trained "global models" for things like channel recommendations and search on customer data, and that users were opted in by default and had to email the company to get out.

That policy sat quietly for months until May 17, 2024, when it went viral on Hacker News (surprise!) and developers started asking why a workplace chat tool needed their conversations to build recommendation models.

A Slack engineer publicly admitted the documentation was outdated, written for an era before Slack AI existed, and needed a rewrite. Slack also clarified that its separate paid AI add-on doesn't train on customer data at all, which raised the obvious question of why the base product's policy read the way it did in the first place.

Adobe had its turn in June 2024 😮‍💨

Its updated terms of service included language letting Adobe access, view, and process user content, including through automated tools, to improve its services with machine learning.

For creative professionals who work under client confidentiality, that read as Adobe reserving the right to look at unreleased client work and feed it into AI training. The backlash was loud enough that Adobe's first clarification blog post didn't settle it, and the company had to promise a further rewrite within days.

It eventually added an explicit line to its terms: "We don't train generative AI on customer content."

Notably, Adobe framed that sentence as a legal obligation it was adding to reassure people, which is a tell in itself. If the concern hadn't been reasonable, there would have been nothing to reassure anyone about.

"To me, this type of surveillance and monetization of young teens’ sense of worthlessness feels like a concrete step toward the dystopian future Facebook’s critics had long warned of."

— Sarah Wynn-Williams, Careless People: A Cautionary Tale of Power, Greed, and Lost Idealism

Grammarly went through its own version in August 2023, when its GrammarlyGO feature launched with a training and retention policy that didn't match what customers were told.

The company's own Chief Information Security Officer (CISO) said Grammarly doesn't let its infrastructure partners use customer data to train their models, and that it kept only anonymized samples for product improvement.

At least one customer reported being told by support that the only way to opt out was to buy a business plan for 500 or more seats, a restriction the CISO later said wasn't accurate. Critics also pointed out that "anonymized" text can still contain passwords and confidential business details, which anonymization doesn't do much to protect.

LinkedIn has since gone through the same motion, opting members into generative AI training on their profile data and activity by default, with an opt-out available for people willing to go find it in settings.

LinkedIn's flow for opting out of the default opt-in data for AI. Source: malwarebytes.com

But wait, there's more! HubSpot gave the pattern its fastest turnaround yet.

On July 1, 2026, the company rolled out a default policy called "Contact Discovery." It opted every CRM customer in automatically, using their data to train AI models and, more strikingly, sharing it with other HubSpot customers.

The backlash was so fast, serve, and loud that HubSpot reversed the policy just five days later, on July 6.

Co-founder and CTO Dharmesh Shah owned it directly on LinkedIn (oh, the irony eh?): "Sorry. You are right. We made a mistake and are reversing that decision," and promised that any future data-sharing feature would be "fully and transparently opt-in," with "clear, upfront control over whether you participate."

Five days is a genuinely fast walk-back by industry standards, and it happened because the backlash was loud immediately, not because the original default was any less deliberate than the others on this list.

What Atlassian's doing fits the same pattern

Which brings us to the story that started this piece. You'd think they'd learn by now, wouldn't you? Nope.

Starting August 17, 2026, Atlassian is using Confluence and Jira content by default to train its Rovo AI, and roughly 300,000 organizations are affected.

On Free and Standard plans, there is no opt-out for the content itself, and metadata like readability scores and task classifications gets collected regardless of tier.

Atlassian's August 17, 2026 data contribution settings update. Source: Atlassian support

Whatever Atlassian keeps can sit in its systems for up to seven years.

Enterprise and Premium customers get a genuine opt-out, and to Atlassian's credit, this policy is published rather than buried in a EULA nobody reads.

But look at what's actually being described. Years of documentation, comments, and tickets, written by a team trying to solve real problems, become training data by default.

Someone has to hold the right subscription tier and go looking for the switch that turns it off.

Line up Zoom, Slack, Adobe, Grammarly, LinkedIn, HubSpot, and Atlassian, and a shape emerges that has nothing to do with any one company being careless.

A vendor ships an AI feature. The feature needs training data, and the company's own customer content is sitting right there, already collected, already paid for by someone else's subscription.

Training-on becomes the default because defaults are where product decisions and legal decisions meet, and the business case for "on" is usually stronger than the business case for "off." The policy gets written down somewhere technically accurate.

Almost nobody reads it until a journalist or a Hacker News thread does, at which point the company clarifies, promises, and sometimes walks part of it back. The default has already moved by then, and moving it back all the way rarely happens.

Seven companies in roughly three years, two of them landing in the same two-month stretch of 2026, is not a coincidence. Every SaaS company is under pressure to ship AI features fast. The cheapest training data any of them will ever find is the content their own paying customers already handed over.

None of this is inevitable

It's worth saying clearly that this isn't the only way to build AI features into a product people already pay for.

Notion's help documentation states its policy without much room for interpretation: "By default, Notion and its AI Subprocessors do not use Customer Data to train any models."

The company backs that with contracts that bind its AI vendors to the same rule, and it's explicit that using Notion AI doesn't hand Notion any license to your content for training purposes. That's a company making the opposite default choice, in writing, and it shows the industry pattern is a choice rather than a law of physics.

Why this matters for a knowledge base

If you run a knowledge base/help center, this question lands differently for you than it does for the average software buyer. A generic wiki page is one thing. A support knowledge base is a detailed, continuously updated record.

It shows exactly how your product breaks, what your customers get confused about, and which workarounds your team quietly relies on. It also shows how your business talks about its own weaknesses when it's trying to help someone.

Internal comments and draft articles often say things out loud that your marketing site never would. That's a strange thing to hand over as training data by default, and it's a stranger thing to find out happened after the fact, buried in a policy page you were never prompted to read.

We built HelpDocs around a simple rule for this. AI features here are opt-in, not opt-out.

If you don't want AI touching your knowledge base, you don't have to go find a setting to turn off, because nothing about it is on until you turn it on. And separately from that choice, we don't train any model on your content.

We think what AI does inside your account should be your call, and whatever you decide, your docs stay yours either way.

We're not going to pretend this is the only sensible way to run a business.

A shared pool of training data can make an AI feature genuinely better over time, faster suggestions, sharper search, and some teams will look at that trade and decide it's worth taking.

That's a real decision worth making on purpose, with the terms in front of you, not one that ships turned on while everyone's attention is on the rest of the changelog.

PDFs walked
so your docs could run.

The modern home for everything your team and customers actually need to find.

HelpDocs platform screenshot