Google’s John Mueller answered a question about robots.txt and explained an easy-to-miss mistake that can impact your SEO and website indexing goals. The specific issue was related to search box spam getting indexed by Google, but this mistake can happen to anyone in general under any context.
Website Search Box Spam
The person who asked the question on Reddit was suffering from a search bar spam attack. What spammers do is search with a query that reflects their spammy niche, and they add a link or a website name. What happens next is that the search bar generates a URL that can be referenced to generate the spammy search result.
And that’s what was happening to the person who was asking the question. Their response was to add a line in the robots.txt file to prevent Google from indexing the file. But Google was indexing those spammy search-generated URLs anyway.
Google Indexed Pages Blocked By Robots.txt
Someone posted on Reddit that their client’s Shopify search box was generating spammy web pages in response to spammer queries and that Google was indexing them despite a robots.txt file prohibiting Google from indexing those pages. What the client did was redirect those spammy URLs to another web page. The person asking the question didn’t ask how to stop the pages from being indexed (which is what they should have been asking); they asked if those redirected URLs should be marked 404 instead.
The person asked:
“Working on a client’s Shopify store where we have the /search added as a disallow in robots.txt, however these search results are still indexed inside of Google.
However, if I try to open one of these pages, they have a redirect set and they redirect to another collection page on the store. Should we display a 404 page instead? What is the easiest way to fix this?”
Why Robots.txt File Caused Spam To Be Indexed
Google’s John Mueller took the extra step to identify and review the client’s robots.txt file and identified an error that was causing Google to ignore the directive prohibiting Googlebot from indexing search results pages.
Mueller responded:
“Also, not sure if it’s your site, but the one I found with similar indexed URLs had sections for “user-agent: Googlebot” (in the “START: Custom Rules” block in comments) as well as a lot more in the “user-agent: *” section further down. With robots.txt, the more specific rules win, so if you have a user-agent: Googlebot section, it will *only* use that section. If you want to apply all the rules in the “user-agent: *” section, you need to copy them. Also, if that’s your site, then you can just list all the user-agents that you want to have shared rules for together, eg:
user-agent: googlebot
user-agent: otherbot
user-agent: imgsrc
user-agent: somethingpt
disallow: /fishes
disallow: /orange-cats
… etc …”
User-Agent Specific Directives Take Precedence
What happened is that the client was relying on Google to follow the directives in a line…
Source link
Disclaimer
We strive to uphold the highest ethical standards in all of our reporting and coverage. We blogs.grocliq.com want to be transparent with our readers about any potential conflicts of interest that may arise in our work. It’s possible that some of the investors we feature may have connections to other businesses, including competitors or companies we write about. However, we want to assure our readers that this will not have any impact on the integrity or impartiality of our reporting. We are committed to delivering accurate, unbiased news and information to our audience, and we will continue to uphold our ethics and principles in all of our work. Thank you for your trust and support.
Website Upgradation is going on for any glitch kindly connect at [email protected]