爬虫遇到取到网页为reload的问题

2022-10-11 12:20:57

有的网站防采集，会在页面加上this.window.location.reload(),这时候你就会得到如下代码：

<html>
   <head>
      <meta http-equiv="Content-Type" content="text/html; charset=UTF-8">
   </head>
   <body>
      <iframe height="0" width="0" style="border: 0px;" src="http://www.***.cn/***/***_cookie.html"></iframe>
      <script type="text/javascript">
setTimeout(function(){
         this.window.location.reload();
                }, 1000);
</script></body>
</html>