Your SaaS subscription has another cost: your data

There is another reason businesses are going to start cutting SaaS products from their stack.
It is not the subscription price.
It is the data.
For years, the trade was pretty straightforward. You gave a software company your customer records, documents, conversations, analytics, sales pipeline, support tickets, internal processes - whatever the product needed - on the assumption that it would be stored safely, locked down, and accessible only to you and the people you explicitly authorised. In return, you got software that would have been far too expensive to build yourself, and you trusted that the provider would keep it secure, private, and properly isolated from everyone else.
That was a reasonable deal.
It isn’t anymore.
A good example happened with HubSpot just a few weeks ago.
HubSpot announced changes to its terms around a new Contact Discovery product and its enrichment dataset. The idea was that customers participating in enrichment could contribute certain professional contact information and email engagement signals to help maintain a shared commercial dataset.
Customers were not impressed.
Four days later, HubSpot reversed the changes.
Their response was unusually direct:
"We made a mistake."
They acknowledged that the rollout made customers feel like the relationship with their CRM, and therefore their data, was changing underneath them. HubSpot cancelled the terms changes and said future enrichment capabilities using customer data would be clearly opt-in.
Credit where it is due. They listened and fixed it.
But the interesting part is that they could propose it in the first place.
Think about what sits inside a CRM.
Your customers.
Your prospects.
Who buys from you.
Who nearly bought from you.
Your notes.
Your sales process.
Years of accumulated commercial intelligence that your team has paid real money to build.
And suddenly you are reading a terms update trying to work out whether some part of that information can contribute to somebody else's commercial dataset.
That would make me uncomfortable too.
One HubSpot customer commenting on the reversal said the software had now been put on their internal "at risk" list. Another said they were initiating an RFP because they no longer felt safe adding commercially valuable data to the platform.
That reaction makes sense to me.
Because once your business runs on somebody else's platform, you do not just depend on their software.
You depend on their future business model.
And AI has made data far more valuable to those business models.
Look at Slack.
Slack says it does not use customer data to train generative AI models unless a customer explicitly opts in.
Good.
But its predictive machine-learning systems can analyse customer data such as messages, content and files to improve global models. If an organisation does not want its data contributing to those models, it has to opt out.
Again, I am not saying Slack is secretly stealing your messages.
That is not the point.
The point is that the data you put into a piece of software can have value to the company running that software beyond simply providing the service you paid for.
And the rules around that can be more complicated than most customers realise.
Google makes the problem even more interesting.
There was a strange Reddit post recently from an indie game developer.
A player was asking Google's Search AI questions about his game. At one point it returned the exact name of an unreleased character: "Vantage Tripod."
The developer said he had never published the name and, as far as he knew, the only digital copy existed inside one of his private Google Docs.
Now, before everybody reaches for the pitchforks, there is no proof that Google leaked his private document.
The name could have come from somewhere else. Metadata. An old build. A forgotten file. Some obscure public source. Nobody has demonstrated what actually happened.
But the story is interesting because Google's own documentation shows how blurry these boundaries are becoming.
For eligible personal Gemini accounts using Connected Apps, Google says data from connected services can be used to improve Google services, including training generative AI models when the relevant activity setting is enabled.
Google says it does not simply dump your entire Gmail inbox or Drive into model training. But summaries, excerpts and inferences from relevant emails and files can be used, and Google explicitly notes that if a file is short or particularly relevant, the "summary" may effectively be the file itself.
Their own documentation tells users not to connect apps containing personal or confidential information they would not want used for generative AI training.
Read that sentence again.
We spent twenty years teaching businesses:
"Put everything in the cloud."
Now we are adding:
"...but make sure you understand which AI settings might cause parts of it to be used to improve somebody else's model."
That is quite a change.
And this is before we get to actual security incidents.
In 2024, attackers compromised Snowflake customer environments using stolen credentials. Mandiant and Snowflake notified around 165 potentially exposed organisations.
Importantly, Mandiant found no evidence that Snowflake's own enterprise environment had been breached. Attackers were getting into customer accounts using credentials stolen elsewhere, helped by things like missing MFA and weak network restrictions.
Technically, that distinction matters.
From the point of view of the company whose data has just been stolen, it probably feels slightly less important.
AT&T disclosed that attackers accessed one of its workspaces on a third-party cloud platform and copied records covering calls and texts for nearly all of its wireless customers during certain periods.
The contents of the calls and texts were not included, nor were things like Social Security numbers or dates of birth.
Still, "nearly all of our wireless customers" is not a sentence anybody wants to put into an SEC filing.
None of this means self-hosting everything is automatically safer.
It is not.
I have seen enough badly configured servers, forgotten admin accounts and applications held together with hope to know that "we built it ourselves" is not a security strategy.
A good SaaS provider can absolutely have better security than a small business.
Often they do.
But there is another side to that equation.
If you use twelve SaaS products, your data now exists across twelve vendors.
Twelve authentication systems.
Twelve sets of employees and contractors.
Twelve lists of subprocessors.
Twelve terms of service.
Twelve companies making their own decisions about AI over the next five years.
And probably a collection of integrations copying data between all of them.
Every new system becomes another place something can go wrong.
That used to be a risk businesses mostly had to accept, because the alternative was spending $100,000 building internal software.
That is the bit AI is changing.
If you are paying $800 a month for a system because you need a database, four screens, a few automations and an email notification, building your own version of those four screens is no longer necessarily insane.
You do not need to recreate Salesforce.
You need to recreate the tiny bit of Salesforce your company actually uses.
You do not need to build Google Workspace.
You might need a small internal application that stores one category of commercially sensitive information somewhere you control.
You do not even necessarily have to "leave the cloud."
There is a huge difference between renting infrastructure from a cloud provider and handing your entire business process to an off-the-shelf SaaS product.
With a purpose-built application, you decide what gets stored.
You decide which external services see it.
You decide whether it is sent to an AI model.
You decide how long it is retained.
And if you want to change those rules, you do not have to wait for the vendor to update its privacy policy.
This is why I think the SaaS argument is becoming much bigger than price.
The monthly saving is nice.
Getting rid of ten features nobody uses is nice.
Having software that actually matches the business is nice.
But owning the rules around your own data might end up being the bigger advantage.
The old calculation was:
Why would we build this when we can rent it for $500 a month?
The new calculation is starting to look more like:
Why are we paying $500 a month to give another company our data, accept their roadmap, accept their security model, accept their future AI policy and use 15% of their product... when building the 15% we actually need has become dramatically easier?
That does not mean every SaaS product should disappear.
I am still not building my own Stripe.
But there are a lot of other subscriptions I would be looking at very closely right now.
And not just because of the invoice.
