Python 并发:asyncio vs ThreadPool vs ProcessPool

Choyeon· 2026年9月5日· 2 分钟阅读· 516 阅读· 508 字· 1,785 字符· 更新于 2026年10月1日
Python 并发:asyncio vs ThreadPool vs ProcessPool

由于全局解释器锁(GIL)的存在,Python 并发性选择常让人困惑。IO 密集和 CPU 密集场景需要完全不同的模型,选错会让性能不升反降。

三大模型原理与代码

线程池提交阻塞函数到系统线程,GIL 在 IO 等待时释放,并发 10-50;进程池子进程彻底绕 GIL,适合 CPU 任务但传参有序列化开销;asyncio 单线程调度协程,适合超大量轻量 IO。

import asyncio, time, requests, httpx
from concurrent.futures import ThreadPoolExecutor, ProcessPoolExecutor, as_completed
from typing import List

URLS = [f"https://jsonplaceholder.typicode.com/posts/{i}" for i in range(1, 51)]

def fetch_sync(url: str) -> dict:
    return requests.get(url, timeout=10).json()

async def fetch_async(client, url: str) -> dict:
    r = await client.get(url, timeout=10.0)
    return r.json()

def run_threadpool(urls, maxw=20):
    t0 = time.perf_counter()
    with ThreadPoolExecutor(max_workers=maxw) as ex:
        fs = [ex.submit(fetch_sync, u) for u in urls]
        results = [f.result() for f in as_completed(fs)]
    print(f"[ThreadPool] {len(results)} tasks, {time.perf_counter()-t0:.2f}s")

def cpu_heavy(n):
    sieve = [True]*(n+1); sieve[0]=sieve[1]=False
    for i in range(2, int(n**0.5)+1):
        if sieve[i]: sieve[i*i::i] = [False]*len(sieve[i*i::i])
    return sum(sieve)

def run_processpool(args):
    t0 = time.perf_counter()
    with ProcessPoolExecutor() as ex:
        results = list(ex.map(cpu_heavy, args))
    print(f"[ProcessPool] sum={sum(results)}, {time.perf_counter()-t0:.2f}s")

async def run_asyncio(urls):
    t0 = time.perf_counter()
    async with httpx.AsyncClient() as client:
        tasks = [fetch_async(client, u) for u in urls]
        results = await asyncio.gather(*tasks)
    print(f"[asyncio] {len(results)} tasks, {time.perf_counter()-t0:.2f}s")

if __name__ == "__main__":
    run_threadpool(URLS)
    run_processpool([100000]*8)
    asyncio.run(run_asyncio(URLS))

选型决策与性能对比

纯 CPU 运算 → ProcessPool;网络/文件 IO → asyncio 或 ThreadPool;混合场景 → asyncio.run_in_executor 把 CPU 任务扔进程池,IO 走协程。

对比维度 ThreadPool ProcessPool asyncio
原理 多线程共享GIL 多进程独立GIL 单线程事件循环+协程
GIL影响 IO时释放/CPU无效 完全无 单线程无竞争
50HTTP请求 ~5-8s ~40s(启动慢) ~1.5-3s
8个CPU素数 ~45s(串行) ~7s(多核) ~40s(单线程)
最大并发规模 数百 CPU核心×2 数千~数万
改造成本 低 中(可pickle) 高(全链路async)

最佳实践

IO 优先 asyncio,第三方库全同步才退用 ThreadPool。CPU 直接多进程,进程数不超过核心数 1-2 倍。

本文作者

评论 (0)

暂无评论,来抢沙发吧。