Understanding Your Data
Suspicious bots
Clearly automated traffic that nobody can vouch for. The honest middle of the classification.
Suspicious is the group for traffic that is obviously a machine, but that we cannot confirm one way or the other. It is not an accusation. It is Honeylog admitting the limit of what can be proven about a request, instead of guessing and presenting the guess as a fact.
It is usually the biggest bot group on a site, and that is normal.
#Why a bot ends up here
There are four common reasons, and they are very different from each other.
- Nobody published anything to check against. Plenty of real, well behaved crawlers are run by smaller companies and research projects that never published the addresses they use. There is simply nothing to compare the request with, so the claim stays unconfirmed.
- It identifies itself as a tool rather than a product. A great deal of automated traffic announces the library or command line tool that made the request instead of a purpose. That is honest, in its way, but it tells you nothing about who is driving it or why.
- It gives no useful identity at all. Requests that are recognisably automated but nameless are grouped together as unknown or generic bots. On most sites this is a large, steady background hum.
- A check could not reach a conclusion. Occasionally a well known crawler cannot be confirmed at the moment its request arrives. We would rather leave that request in the middle than accuse a real search engine of being an impostor, so it lands here too.
#How to read this group
Volume alone means very little here. What is worth your attention is shape and change.
A monitoring service checking one page every minute is a flat line and a non event. A nameless bot that appears from nowhere and works methodically through your entire product catalogue is a scraper, whatever it calls itself, and the pattern is visible long before anybody can prove ownership. The same is true of automated traffic that suddenly concentrates on your pricing page, your search results, or anything behind a form.
So the useful questions are which pages it is reaching, whether the volume is steady or climbing, and whether it showed up recently. The report groups by page and by name, and you can set an alert if you want to be told when this traffic grows rather than going to look.
#When suspicious becomes something else
This group is not a permanent verdict. As the bot catalogue grows and more organisations publish what their crawlers use, traffic that sits here today can be recognised tomorrow, and future requests from that crawler will be verified instead. Requests already recorded keep the verdict they were given at the time, which is why an old report and a new one can describe the same crawler differently.