
You probably put a third-party traffic chart in a client deck this month. A Similarweb screenshot showing a competitor’s monthly visits, a Semrush visibility graph, an Ahrefs backlink count. I have done it hundreds of times.
Here is a question I had never asked in 25 years of doing this work: does the license actually let me?
The question came up at BrightonSEO in San Diego a few days ago, in a conversation at a vendor booth about whether their data belongs in a legal case. I went back to my desk and read thirty license agreements covering the tools our industry runs on. Search visibility platforms, crawlers, link indexes, every major ad platform, analytics, panel measurement, domain forensics. All of them read in full, all of them coded the same way.
What I found matters more for the chart in your deck than for anything a court will ever see. I published the full findings as a 28-page research report, and this is the part digital marketers need.
Ten of Thirty Require Permission Before You Publish
The clause that should worry you is not about lawyers. It is about presentation.
Section 6(iv) of the Similarweb terms of service, effective March 19, 2025, lists among the prohibited acts:
present or share the data or information received through the Platform without Similarweb’s prior consent, and in the event such consent was given, present or share such data or information without attribution to Similarweb pursuant to Similarweb’s branding guidelines
Read it as a marketer rather than as a litigator. A blog post is presenting data received through the Platform. So is a conference slide. So is a webinar, a pitch deck, a LinkedIn carousel, and a client proposal that gets forwarded to three people you have never met.
Ten of the thirty agreements I read require prior written consent to publish, or prohibit publication outright. SpyFu bars distribution without written consent, with a narrow exception that requires the phrase “Used with permission from SpyFu.com” on anything you do share. SparkToro prohibits reselling or white-labeling its output, then carves out an exception letting agencies use reasonable excerpts in client deliverables with attribution. Ahrefs bars publishing any portion of its services for commercial purposes, language written with resale in mind but broad enough that I would ask first.
Twelve of the thirty impose no publication restriction at all. Semrush is the most permissive of the visibility tools and bars sublicensing, scraping and AI training without restricting a chart. Your own Google Analytics and Search Console data carries no restriction whatsoever, because it is yours.
The practical fix takes one email. Vendors say yes, because a cited chart is free marketing. I have never had one refuse.
Your Ad Performance Data May Not Be Yours
This is the finding that surprised me most, and it affects every paid media team.
Amazon’s advertising agreement, posted July 23, 2026, states that all ad services data is Amazon’s exclusive property, restricts what you may do with it, and permits disclosure only where a valid and binding court order requires it, on condition that you tell Amazon before you disclose. LinkedIn’s ads agreement, updated November 3, 2025, expressly designates its metrics and professional demographics as confidential information and limits your use of ad services data to aggregate and anonymous performance assessment. TikTok’s commercial terms restrict sharing platform data with third parties.
Microsoft Advertising takes the opposite approach and includes no confidentiality clause at all. Meta does not designate campaign data as confidential, and publishes political advertisers’ spend, reach and targeting in its Ad Library for seven years.
So two of the four largest ad platforms in the world claim your campaign performance data and two do not. Nobody mentions which is which when you open the account. If you publish client benchmarks, run a public performance study, or put spend figures in a case study, that distinction decides whether you are inside the agreement.
The Accuracy Gap Nobody Talks About
Twenty-seven of the thirty agreements disclaim the accuracy of their data. Three describe that data as estimates or modeled output anywhere in their published terms.
Those are not the same thing, and the difference is the whole argument. A warranty disclaimer is a liability instrument. It says the vendor will not be answerable if you rely on the number and lose money. A statement of measurement uncertainty is a methodological instrument. It tells you what the number is, how it was produced, and how far it is likely to sit from the thing it estimates.
Similarweb states that its data rests on estimations and extrapolations. Google publishes a caveat that Trends data should be treated as one data point among others and is not a perfect mirror of search activity. Nielsen publishes real sampling error documentation. The other twenty-seven give you a disclaimer and nothing else. Majestic comes closest to candor without using the word, explaining that the shifting structure of the web makes it “impossible to provide completely accurate link data.”
How Far Off Are the Estimates? There Is an Actual Answer
Our industry argues about this in blog comments. A peer-reviewed study settled it with data.
Jansen, Jung and Salminen, publishing in PLOS ONE in 2022, compared Similarweb against the sites’ own Google Analytics across 86 websites in 26 countries and 19 industry verticals over twelve months. Similarweb understated total visits by 19.4 percent and unique visitors by 38.7 percent. It overstated bounce rate by 25.2 percent and average session duration by 56.2 percent.
The number that matters is not the size of the gap. It is the shape. The bias runs in a consistent direction, which means these estimates are not noisy, they are skewed, and a skew you can name is a skew you can account for. Traffic reads low. Engagement reads high.
The problem is not unique to one vendor. Research at the ACM Internet Measurement Conference in 2022 checked popularity rankings against Cloudflare server-side request data and found 87.1 percent of the Alexa top thousand were overranked, with 56.7 percent overranked by two or more orders of magnitude.
None of that makes the tools useless. It makes them what they have always been: modeled approximations that are excellent for direction and pattern and unfit to carry a precise factual claim on their own. Say “Similarweb estimates roughly 2 million visits” and you are describing the instrument correctly. Say “they get 2.1 million visits a month” and you have made a statement the vendor declined to make.
Where This Gets Expensive
The legal side is my day job, so one paragraph on it. A terms of service is a contract between you and a vendor. It cannot make a document inadmissible, and it is not a defense to a subpoena. That is governed by the rules of evidence and decided by a judge. What a restrictive license can do is put you in breach, which costs you the account or a damages claim, and it can collide with the disclosure obligations that apply when you are retained as an expert witness. Exactly one of the thirty agreements addresses litigation at all, and that one prohibits it.
The report covers all of that in detail. For most readers here, the courtroom is not the risk. The risk is publishing something your license did not permit, or publishing an estimate as though it were a measurement, in front of an audience that includes your next client.
What I Would Do Monday Morning
Read the licenses for the tools you actually publish from. It took me an afternoon for thirty of them, and I had been paying for some for a decade without opening the agreement once.
Ask permission where the terms require it, and keep the reply. Follow the attribution format the vendor specifies, because several of them dictate the wording.
Label estimates as estimates, with the tool, the metric, the date range and the known limitation in the same paragraph as the figure. A footnote is not a disclosure. When I analyzed twenty years of my own backlinks, I published the index, the export, the statistical method and the limitation that biases the result, and the limitation was the part readers wrote to me about. Disclosure did not weaken the analysis. It was the reason anyone trusted it.
Apply the same skepticism to other people’s numbers. I took apart a Semrush study of ChatGPT categories not because the data was bad but because the headline number everyone repeated described something other than what readers assumed.
And build the important work on data you own. Your analytics, your Search Console, your ad account exports, your server logs, your own crawl. Nobody licenses those to you and nobody can revoke them. That is true whether you are diagnosing a traffic and rankings loss or writing the study that wins you a speaking slot.
The tools are good. The claims we make about them are often worse than the tools deserve. Every platform in this survey measures something real by a method with known limits, and those limits are not a defect. They are the specification.
The full research report is free. Thirty license agreements, the coding for each, the evidence rules, the reliability research and the complete source list. Download the 28-page PDF here.