首页 > 网站技巧 > 服务器 > Linux > apache禁止收录爬虫采集

apache禁止搜索引擎收录、网络爬虫采集的配置方法

2014-06-30 08:48:10 投稿：junjie

这篇文章主要介绍了apache禁止搜索引擎收录、网络爬虫采集的配置方法,注意一定要写到Location节点,否则不起作用,可以精确匹配,也可以IP匹配,需要的朋友可以参考下

Apache中禁止网络爬虫，之前设置了很多次的，但总是不起作用，原来是是写错了，不能写到Dirctory中，要写到Location中

<Location />

SetEnvIfNoCase User-Agent "spider" bad_bot

BrowserMatchNoCase bingbot bad_bot

BrowserMatchNoCase Googlebot bad_bot

Order Deny,Allow

#下面是禁止soso的爬虫

Deny from 124.115.4. 124.115.0. 64.69.34.135 216.240.136.125 218.15.197.69 155.69.160.99 58.60.13. 121.14.96. 58.60.14. 58.61.164. 202.108.7.209

Deny from env=bad_bot

</Location>

这是禁止了所有包含spider字符的爬虫。
如果要针对性的禁止爬虫，改成精确匹配的爬虫字符串，如果bingbot、Googlebot等等

apache禁止搜索引擎收录、网络爬虫采集的配置方法

您可能感兴趣的文章: