互联网 qkzz.net
全刊杂志网:首页 > 女性 > 文章正文
刊社推荐

基于GNP算法的分布式爬虫调度策略刘 爽 姜春祥 张伟哲 李 东 张 鸿


摘 要:针对分布式搜索引擎的任务调度及负载均衡问题,提出了基于GNP算法的分布式爬虫调度策略和负载均衡的方法。利用网络距离预估取代大规模的网络距离测量,不仅提高了系统的响应速度,还减少了系统对广域网造成的压力。通过在广域网上部署爬虫节点,构建分布式搜索引擎,应用该调度策略进行实验,验证了系统性能有较大提高。
  关键词:分布式爬虫; 任务调度; 负载均衡; 网络测量; 全局网络定位
  中图分类号:TP309
  文献标志码:A
  
  文章编号:1001-3695(2010)02-0446-04
  doi:10.3969/j.issn.1001-3695.2010.02.011
  
  GNP-based scheduling strategy for distributed crawling
  
  LIU Shuang1, JIANG Chun-xiang2, ZHANG Wei-zhe1, LI Dong1, ZHANG Hong3
  
  (1.School of Computer Science & Technology, Harbin Institute of Technology, Harbin 150001, China; 2.Heilongjiang Branch of National Computer Network Emergency Response Technical Team/Coordination center of China, Harbin 150001, China; 3.National Computer Network Emergency Response Technical Team/Coordination Center of China, Beijing 100029, China)
  
  Abstract:In order to solve task scheduling and load balancing problems of distributed search engines, this paper proposed a GNP-based scheduling strategy for distributed crawling and a load balancing method. Adopted internet distance estimating mechanism as a replacement for large-scale network distance measurement, which not only improved response time of the system, but also reduced WAN pressure caused by the system. Through deploying crawling nodes at WANs, built a distributed search engine, and implemented several scheduling strategies. The online experiment shows great improvement in system’s performance. ......
很抱歉,暂无全文,若需要阅读全文或喜欢本刊物请联系《计算机应用研究》杂志社购买。
欢迎作者提供全文,请点击编辑
分享:
 

了解更多资讯,请关注“木兰百花园”
分享:
 
精彩图文


关键字
支持中国杂志产业发展,请购买、订阅纸质杂志,欢迎杂志社提供过刊、样刊及电子版。
关于我们 | 网站声明 | 刊社管理 | 网站地图 | 联系方式 | 中图分类法 | RSS 2.0订阅 | IP查询
全刊杂志赏析网 2017