一、问题背景:为什么向量库选型会卡住项目

我们做的是一个图文内容推荐系统,每天新增约80万条内容向量,模型输出768维,历史存量约1000万条,未来半年预计到5000万。线上检索要求:

  • 单次查询返回Top 50,带元数据过滤(类目、时间、状态)
  • 峰值QPS 200,P99 < 50ms
  • 单节点可落地,后续能平滑扩到集群
  • 运维成本尽量低,团队没有专职向量数据库DBA

候选就是Milvus和Qdrant。网上文章要么只讲概念,要么只跑10万条数据,没有参考价值。我干脆在同一台机器上把两个都部署了,用真实数据和真实查询模式压了一遍。

二、环境与版本

  • 机器:AWS EC2 c6i.2xlarge,8 vCPU,32GB内存,GP3 500GB
  • OS:Ubuntu 22.04 LTS,内核5.15
  • Docker:24.0.7,Docker Compose v2.24.5
  • Milvus:2.4.1(standalone模式,etcd 3.5.5 + MinIO RELEASE.2023-03-20)
  • Qdrant:1.9.0(单节点)
  • 客户端:Python 3.11,pymilvus 2.4.3,qdrant-client 1.9.1
  • 数据:1000万条768维float32向量,随机生成但分布接近真实embedding(归一化后)

三、方案设计

两条路线保持尽可能公平:

  1. 都使用HNSW索引。Milvus参数:M=16, efConstruction=200, efSearch=64;Qdrant参数:m=16, ef_construct=200, hnsw_ef=64。
  2. 距离度量统一用COSINE。
  3. 元数据过滤字段一致:category(int)、ts(int)、status(int)。
  4. 压测工具用locust + 自写Python脚本,查询向量从测试集中随机取,避免缓存全命中。
  5. 资源监控用docker stats和pidstat,每10秒采样。

四、核心实现

4.1 部署步骤

Milvus standalone的compose文件官方就有,我精简后如下:

# docker-compose-milvus.yml
version: '3.5'
services:
  etcd:
    image: quay.io/coreos/etcd:v3.5.5
    environment:
      - ETCD_AUTO_COMPACTION_MODE=revision
      - ETCD_AUTO_COMPACTION_RETENTION=1000
      - ETCD_QUOTA_BACKEND_BYTES=4294967296
    volumes:
      - ./volumes/etcd:/etcd
    command: etcd -advertise-client-urls=http://127.0.0.1:2379 -listen-client-urls http://0.0.0.0:2379 --data-dir /etcd
    healthcheck:
      test: ["CMD", "etcdctl", "endpoint", "health"]
      interval: 30s
      timeout: 20s
      retries: 3

  minio:
    image: minio/minio:RELEASE.2023-03-20T20-16-18Z
    environment:
      MINIO_ACCESS_KEY: minioadmin
      MINIO_SECRET_KEY: minioadmin
    volumes:
      - ./volumes/minio:/minio_data
    command: minio server /minio_data --console-address ":9001"
    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:9000/minio/health/live"]
      interval: 30s
      timeout: 20s
      retries: 3

  standalone:
    image: milvusdb/milvus:v2.4.1
    command: ["milvus", "run", "standalone"]
    environment:
      ETCD_ENDPOINTS: etcd:2379
      MINIO_ADDRESS: minio:9000
    volumes:
      - ./volumes/milvus:/var/lib/milvus
    ports:
      - "19530:19530"
      - "9091:9091"
    depends_on:
      - etcd
      - minio

启动:

docker compose -f docker-compose-milvus.yml up -d
docker compose -f docker-compose-milvus.yml ps

Qdrant就简单得多,单容器:

docker run -d --name qdrant \
  -p 6333:6333 -p 6334:6334 \
  -v $(pwd)/qdrant_storage:/qdrant/storage \
  qdrant/qdrant:v1.9.0

4.2 写入与索引代码

Milvus写入:

from pymilvus import connections, Collection, CollectionSchema, FieldSchema, DataType, utility
import numpy as np

connections.connect("default", host="localhost", port="19530")

fields = [
    FieldSchema(name="id", dtype=DataType.INT64, is_primary=True),
    FieldSchema(name="vec", dtype=DataType.FLOAT_VECTOR, dim=768),
    FieldSchema(name="category", dtype=DataType.INT32),
    FieldSchema(name="ts", dtype=DataType.INT64),
    FieldSchema(name="status", dtype=DataType.INT32),
]
schema = CollectionSchema(fields, description="content_embedding")
col = Collection("content_milvus", schema, consistency_level="Bounded")

index_params = {
    "index_type": "HNSW",
    "metric_type": "COSINE",
    "params": {"M": 16, "efConstruction": 200}
}
col.create_index("vec", index_params)
col.load()

# 分批写入,每批5万
BATCH = 50000
for i in range(0, 10_000_000, BATCH):
    n = min(BATCH, 10_000_000 - i)
    ids = list(range(i, i + n))
    vecs = np.random.rand(n, 768).astype(np.float32)
    cats = np.random.randint(0, 20, n).tolist()
    tss = np.random.randint(1700000000, 1720000000, n).tolist()
    sts = np.random.randint(0, 2, n).tolist()
    col.insert([ids, vecs, cats, tss, sts])
col.flush()

Qdrant写入:

from qdrant_client import QdrantClient
from qdrant_client.models import VectorParams, Distance, PointStruct, HnswConfigDiff
import numpy as np

client = QdrantClient(host="localhost", port=6333)

client.recreate_collection(
    collection_name="content_qdrant",
    vectors_config=VectorParams(size=768, distance=Distance.COSINE),
    hnsw_config=HnswConfigDiff(m=16, ef_construct=200),
)

BATCH = 5000
points = []
for i in range(10_000_000):
    points.append(PointStruct(
        id=i,
        vector=np.random.rand(768).astype(np.float32).tolist(),
        payload={
            "category": int(np.random.randint(0, 20)),
            "ts": int(np.random.randint(1700000000, 1720000000)),
            "status": int(np.random.randint(0, 2)),
        }
    ))
    if len(points) == BATCH:
        client.upsert(collection_name="content_qdrant", points=points, wait=False)
        points = []
if points:
    client.upsert(collection_name="content_qdrant", points=points, wait=True)

注意Qdrant的wait=False在批量写入时非常关键,默认同步等待会慢一个数量级。

五、踩坑与优化

坑1:Milvus standalone内存爆掉。 默认配置下,Milvus的queryNode和dataNode会吃掉大量内存,1000万条768维向量建HNSW索引时,32G机器直接OOM。解决办法是在milvus.yaml里限制:

queryNode:
  cache:
    cacheSize: 8GB
dataNode:
  memory:
    forceSyncEnable: true

或者用环境变量MILVUS_QUERYNODE_CACHE_CACHESIZE=8GB。改完重建容器,索引阶段峰值内存从31G降到22G。

坑2:Qdrant的payload索引没建,过滤查询慢10倍。 一开始只建了向量索引,带category=3 AND status=1的查询P99到了180ms。后来对过滤字段建payload索引:

client.create_payload_index("content_qdrant", "category", field_schema="integer")
client.create_payload_index("content_qdrant", "status", field_schema="integer")

重建后同样查询P99降到38ms。

坑3:Milvus的consistency_level选错。 默认Strong会导致每次查询都等flush,QPS上不去。改成Bounded后,写入后1秒内可见,查询QPS提升约35%。业务上我们能接受1秒延迟。

坑4:Qdrant的hnsw_ef是查询时参数。 建索引时的ef_construct和查询时的hnsw_ef是两回事,一开始只设了前者,召回率只有0.82。查询时传search_params={"hnsw_ef": 64}后,召回率到0.96。

六、效果数据

压测持续30分钟,每轮5分钟,取稳定段数据。

指标 Milvus 2.4.1 Qdrant 1.9.0
索引构建时间 42分钟 28分钟
索引后磁盘占用 38GB 31GB
空闲内存占用 6.2GB 1.8GB
压测峰值内存 22.4GB 9.6GB
QPS(无过滤) 312 486
QPS(带过滤) 187 341
P99(无过滤) 41ms 26ms
P99(带过滤) 68ms 38ms
召回率@50 0.97 0.96
冷启动加载时间 3分20秒 48秒

写入方面,Milvus批量写入1000万条耗时约19分钟,Qdrant约14分钟,但Qdrant在wait=True时单条写入延迟明显高于Milvus,批量场景两者都能接受。

七、总结

最终我们选了Qdrant。原因很直接:单节点资源占用低、部署简单、带过滤查询性能更好,完全满足当前1000万级、QPS 200的需求,而且运维成本几乎为零。Milvus不是不好,它的分布式能力、多索引类型、生态成熟度在5000万以上或需要多租户隔离时优势明显,但对我们这个阶段来说太重了。

选型建议给三条:

  1. 单节点、数据量5000万以下、过滤查询多,优先Qdrant。
  2. 需要分布式、多副本、强一致、或者已经在用Milvus生态,选Milvus。
  3. 无论选哪个,一定要用真实数据和真实查询模式压测,别信benchmark网站的数字,你的过滤条件和召回率要求会彻底改变结论。

代码和compose文件我都放在仓库里了,改改就能跑。选型这事,跑一遍比看十篇文章都管用。