This Summer has been “interesting” in the world of data collection from the web and, especially, from Google. Long story short, Google does not want the AI labs to use their SERPs as part of their live retrieval process to answer questions. Everyone will have their own opinions on whether or not that is a good thing but this is what we believe is happening.
I’ve been working in this space for 20+ years and the stretch since early September has been the worst that we have seen in terms of challenges to acquire this data.
Seizure of NetNut’s network
Things were relatively quiet on the data acquisition side of things until early July. You may have seen that the FBI / DOJ appeared to have worked with Google to seize the proxy network of NetNut, one of the largest IP networks in the world, amid allegations that NetNut had built a “botnet,” which is a very serious allegation to make. Google’s Threat Intelligence Group put the figure at around two million devices; the parent company disputes the allegations. Whether or not they were doing such a thing is outside of my purview but I am skeptical. Rather, my guess is that this is simply the framing that was adopted in order to mobilize law enforcement. Again, this is all just my own speculation.
Long story short, this had the effect of removing a large provider from the global network of proxies, which affects everyone in the space. This had an impact on us as well, and so we spent additional time during July routing around this and building new relationships.
Goto links
Things stabilized for awhile after the July festivities in August and then Google deployed another countermeasure, what has come to be known as the Goto links.
This was designed to make it difficult to actually resolve what URLs are ranking or appearing in SERP Features.
We were, fortunately, able to resolve this one very quickly. In a matter of hours. The reason for this one was because we already had the infrastructure in place due to the fact that we needed it for other engines.
The September Ban Hammer
Right around September 11th, Google deployed the most aggressive blocking that we have seen to date. This continued through the weekend and has continued since then. The weird part about it was that nobody was talking about it in public. I’ll talk about this more in the next section.
The blocking operation is still ongoing as of this writing and it represents, I believe, the new reality going forward. Now that it has been a few weeks, the industry is collectively starting to recover and a whole new host of techniques have been used. Operational playbooks that were in place for 5+ years had to be rewritten.
I’m not going to go into much more detail than that and I would recommend that other companies follow suit. This isn’t about us and our operations, I just have come to the conclusion that it can cause more damage to our customers to share too much information about specific remedies in public.
Deep conversations continue behind closed doors and we are speaking regularly with many of the top providers and other vendors in this space.
An eerie silence
The September incident was interesting in our community because the vibe shifted.
We started to notice that everyone became a little more quiet, including ourselves, about the fact that there even was an issue. Typically, when there is a large outage, it ends up on the front page of Search Engine Land and everybody starts asking (as they have been for 20 years) if “the rank trackers are finally dead.” (As always, the answer is no.)
We’ve also been much more circumspect about the countermeasures we have deployed. The reason for this is, ultimately, we’re living in a world now of information asymmetry between Google and everyone trying to get access to their data on Google.
It doesn’t make sense for people to be as open as they used to be about the workarounds and techniques to make it possible. We deliberately chose to be vague in our messaging as well because, if it becomes more difficult to acquire this data, it hurts our customers.
As a result, DemandSphere, along with many of the other vendors in our space, has started to take the approach of keeping the public messaging somewhat vague, with more direct and open communication to customers.
We do continue to provide updates about everything possible on our Status Dashboard and encourage all of our customers to subscribe to notifications there.
We think we know what they will do next
Unfortunately, we don’t think they are done with making it more difficult. We have a pretty good idea of the sorts of things they will do next and, indeed, are already starting to see hints of it.
We’re not going to share too much about this, for the reasons mentioned above, in public forums but will be having ongoing conversations with our clients and the community.
What they should do instead
Innovation is always better than playing defense, in my opinion. When you build a business that essentially achieves utility scale in the world, both in terms of size and impact, you’re looking at potential antitrust questions from every angle.
I understand that this is a tightrope that they are walking but their best bet is to continue to create the best user experiences and just assume that this data is going to be collected, one way or another. Because it will be.
What this is about
Google doesn’t care about tools in the AI search / SEO space. Companies in our industry are barely a blip on their radar. The activities of the other AI companies (OpenAI, Anthropic, etc.) are what worry them.
They view these companies as stealing their data and repurposing it for their own profits, while not giving proper traffic and attribution to the source.
Ironies abound.
Why we believe that this data is yours
Our view is simple: if there is data about you, or data that affects you, it belongs to you.
If the originators of that data won’t find a way to provide it, even on reasonable commercial terms, then a market will exist to provide it.
All of Google’s legal challenges to this have fallen flat, which is why we believe they are taking more aggressive means to prevent it from happening.
Ultimately, we don’t think their efforts will matter much in the end. It will become more difficult, and we already have an idea of what they plan to do next. It will be annoying for awhile and there will be more outages (and other things), which we log alongside every confirmed Google update but, because this data is so important, the market will find a way to provide what people need.
As I mentioned above, I don’t think this is targeted at the SEO platforms. Frankly, I think we’re all too small, including even the biggest platforms in our category, to have much of an impact. I think this is about the AI labs, and everybody else is just caught in the crossfire.