PerplexityBot爆发式访问不厚道,还不支持Crawl-delay!(吐槽)
发现国外这个PerplexityBot爬虫有点过分,访问是爆发式的,而且根据网上消息还不支持Crawl-delay值的设置!
访问记录如下(摘):
1 2026-06-15 02:51:09 GET /*/*.html - 80 - 18.97.9.102 Mozilla/5.0+AppleWebKit/537.36+(KHTML,+like+Gecko;+compatible;+PerplexityBot/1.0;++https //perplexity ai/perplexitybot) - 200 0 0 2 2026-06-15 02:51:09 GET /*/*.html - 80 - 18.97.9.101 Mozilla/5.0+AppleWebKit/537.36+(KHTML,+like+Gecko;+compatible;+PerplexityBot/1.0;++https //perplexity ai/perplexitybot) - 200 0 0 3 2026-06-15 02:51:09 GET /*/*.html - 80 - 18.97.9.101 Mozilla/5.0+AppleWebKit/537.36+(KHTML,+like+Gecko;+compatible;+PerplexityBot/1.0;++https //perplexity ai/perplexitybot) - 200 0 0 4 2026-06-15 02:51:09 GET /*/*.html - 80 - 18.97.9.96 Mozilla/5.0+AppleWebKit/537.36+(KHTML,+like+Gecko;+compatible;+PerplexityBot/1.0;++https //perplexity ai/perplexitybot) - 200 0 0 5 2026-06-15 02:51:09 GET /*/*.html - 80 - 18.97.9.99 Mozilla/5.0+AppleWebKit/537.36+(KHTML,+like+Gecko;+compatible;+PerplexityBot/1.0;++https //perplexity ai/perplexitybot) - 200 0 0
ip挺分散的,哪怕Crawl-delay设置了10秒,估计都难以阻挡。以下是PerplexityBot爬虫不支持Crawl-delay的有关证据:
| 序号 | 概述 | 说明 |
| 证据1 | Perplexity 官方自身的声明 | Perplexity 官方在提供给网站管理员的文档中,明确指出其 Perplexity-User 类型的请求(即用户在使用 Perplexity 时发起的实时检索)“generally ignores robots.txt rules”(通常会忽略 robots.txt 规则) |
| 证据2 | Cloudflare 等第三方权威机构的实测报告 |
网络安全公司 Cloudflare 在 2025 年进行了一次严格的实地测试,以验证 Perplexity 对 robots.txt 的遵从情况: 实验设计:Cloudflare 创建了全新的、从未被公开过的测试域名,并设置严格的 robots.txt 规则,明确禁止所有爬虫(包括 PerplexityBot)访问任何内容。 实验结果:当通过 Perplexity AI 搜索这些域名时,它依然能够提供被其爬虫抓取到的详细内容信息。 这项由权威第三方进行的严格测试,提供了直接、公开的证据,证明 Perplexity 的爬虫体系能够无视 robots.txt 的指令。 |
| 证据3 | 使用“隐身”爬虫绕开封锁 |
Cloudflare 报告指出,当 Perplexity 官方的 PerplexityBot 被网络防火墙(如 WAF)或 robots.txt 成功阻挡后,其系统会启动“隐身”爬虫: 伪造身份:它会将用户代理(User-Agent)修改为普通浏览器的样子,例如伪装成 Google Chrome on macOS,以冒充真实用户流量。 隐藏来源:它还会不断变换其 IP 地址,超出其官方公布的范围,使传统的网络封锁失效。 |
| 证据4 | Crawl-delay 本身并非官方标准 | 从技术标准来看,Crawl-delay 指令本身也不是一个万能的解决方案。它并非官方 robots.txt 标准的一部分,因此并非所有爬虫都会支持或遵守它。就连 Googlebot 这样的主流搜索引擎爬虫也明确忽略 Crawl-delay 指令。 |

浙公网安备 33010602011771号