热门问题python爬虫的效率如何提高_Python

简单版本爬虫

我们先来一个简单的爬虫，看看单线程处理会花费多少时间？

				?

									import time

									import requests

									from datetime import datetime

									def fetch(url):

									    r = requests.get(url)

									    print(r.text)

									start = datetime.now() 

									t1 = time.time()

									for i in range(100):

									    fetch('http://httpbin.org/get') 

									print('requests版爬虫耗时：', time.time() - t1)

									# requests版爬虫耗时：54.86306357383728

我们用一个爬虫的测试网站，测试爬取100次，用时是54.86秒。

多线程版本爬虫

下面我们将上面的程序改为多线程版本：

				?

									import threading

									import time

									import requests

									def fetch():

									    r = requests.get('http://httpbin.org/get')

									    print(r.text)

									t1 = time.time()

									t_list = []

									for i in range(100):

									    t = threading.Thread(target=fetch, args=())

									    t_list.append(t)

									    t.start() 

									for t in t_list:

									    t.join() 

									print("多线程版爬虫耗时：", time.time() - t1)

									# 多线程版爬虫耗时：0.8038511276245117

我们可以看到，用上多线程之后，速度提高了68倍。其实用这种方式的话，由于我们并发操作，所以跑100次跟跑一次的时间基本是一致的。这只是一个简单的例子，实际情况中我们不可能无限制地增加线程数。

多进程版本爬虫

除了多线程之外，我们还可以使用多进程来提高爬虫速度：

				?

									import requests

									import time

									import multiprocessing

									from multiprocessing import Pool

									MAX_WORKER_NUM = multiprocessing.cpu_count() 

									def fetch():

									    r = requests.get('http://httpbin.org/get')

									    print(r.text) 

									if __name__ == '__main__':

									    t1 = time.time()

									    p = Pool(MAX_WORKER_NUM)

									    for i in range(100):

									        p.apply_async(fetch, args=())

									    p.close()

									    p.join()

									    print('多进程爬虫耗时：', time.time() - t1)

									多进程爬虫耗时： 7.9846765995025635

我们可以看到多进程处理的时间是多线程的10倍，比单线程版本快7倍。

协程版本爬虫

我们将程序改为使用 aiohttp 来实现，看看效率如何：

				?

									import aiohttp

									import asyncio

									import time 

									async def fetch(client):

									    async with client.get('http://httpbin.org/get') as resp:

									        assert resp.status == 200

									        return await resp.text() 

									async def main():

									    async with aiohttp.ClientSession() as client:

									        html = await fetch(client)

									        print(html) 

									loop = asyncio.get_event_loop() 

									tasks = []

									for i in range(100):

									    task = loop.create_task(main())

									    tasks.append(task) 

									t1 = time.time() 

									loop.run_until_complete(main()) 

									print("aiohttp版爬虫耗时：", time.time() - t1) 

									aiohttp版爬虫耗时： 0.6133313179016113