Understanding Your Data

Fake bots

Requests that claim an identity the evidence contradicts, including software that pretends to be a person.

A fake bot is a request whose story does not hold up. Something about it claims one identity while the evidence points somewhere else. This is the one group Honeylog is willing to call out directly, because in these cases there is no ambiguity left to respect.

#Impersonators

The classic case is a request that says it is Googlebot and arrives from an address Google does not use. Or SemrushBot from an address that has nothing to do with Semrush. The name is being borrowed because well known crawlers get treated well: sites let them through, serve them quickly, and rarely question them.

In your reports these keep the name they claimed, with the word Fake in front. Seeing Fake Googlebot next to Googlebot is deliberate, because the useful thing to know is not only that someone lied, but who they chose to impersonate. A wave of traffic pretending to be an AI crawler means something quite different from a wave pretending to be a search engine.

#Software pretending to be people

The larger and quieter family is traffic that makes no bot claim at all. It presents itself as an ordinary browser, on an ordinary operating system, exactly like a real visitor would, and it shows up in your reports as Fake Human.

What gives it away is where it comes from. Real people browse from home connections, mobile networks, and offices. This traffic arrives from hosting providers, the rented machines that run servers, which is not a place anybody reads a blog post from. A browser signature coming out of a server rack is a script wearing a costume.

They tend to rotate through thousands of different browser signatures to look like a crowd, which is another thing the reports let you see rather than hide.

Alongside them you will sometimes find traffic from addresses already known for abuse, grouped as spam, which is treated the same way for the same reason.

#Why keeping them separate matters

If this traffic were counted as people, your audience numbers would be inflated by machines, and the sites with the worst problem would look like the ones doing best. Honeylog never folds fake traffic into the human figures. It is counted, shown, and named, and it stays out of the numbers you report to other people.

#What you can do about it

Politeness is wasted here. A crawler that ignores your rules while impersonating somebody else is not going to respect a preference file, so the realistic options are at the level where requests are actually served or blocked, in front of your site.

What Honeylog gives you is the evidence: how much of it there is, which pages it is going after, where it comes from, and whether it is growing. You can also set an alert on this traffic so you hear about a surge when it starts rather than at the end of the month.