list_collections 与 count_collections

🎉摘要:本文详细介绍 ChromaDB 向量数据库中如何获取当前租户和数据库下的所有向量集合信息,包括 list_collections 的分页查询方法及高效统计集合数量的 count_collections 函数,并提供 Python 示例代码。

如果你要获取当前租户+数据库下面所有的向量集合信息,可以使用 list_collections 函数。

函数签名如下:

client.list_collections(limit=None, offset=None)

其中:

  • limit  最多返回多少个集合。默认上限 100 条,超过 100 必须分页

  • offset   分页偏移量,从第 offset 条之后开始取  

例如:

import chromadb

client = chromadb.EphemeralClient()

def create(name:str, num:int) -> None:
    col = client.create_collection(name=name, embedding_function=None)
    col.add(
        ids=[f"qa{i}" for i in range(num)],
        documents=[f"这是问题{i}" for i in range(num)],
        metadatas=[{"tag":f"tag{i}"} for i in range(num)]
    )

# 准备数据
create("my_col1", 2)
create("my_col2", 10)
create("my_col3", 5)

# 查询集合信息
cols = client.list_collections()
for c in cols:
    print(f"集合名:{c.name}, 数量 {c.count()}")

# 输出:
# 集合名:my_col3, 数量 5
# 集合名:my_col2, 数量 10
# 集合名:my_col1, 数量 2

注意,list_collections() 返回的是 Collection 对象列表,按创建时间从旧到新排序。

默认最多返回 100 个。超过要用分页:

# 前 100 个
batch1 = client.list_collections(limit=100, offset=0)

# 后 100 个
batch2 = client.list_collections(limit=100, offset=100)

# 从第 50 个开始取 20 个
subset = client.list_collections(limit=20, offset=50)

如果你只想要数量,请使用 count_collections() 函数。 统计当前 database 内一共有多少个集合,直接返回整数,不需要遍历全部集合,性能远优于 len(client.list_collections())。例如:

import chromadb

client = chromadb.EphemeralClient()
# 准备数据
client.create_collection("my_col1", embedding_function=None)
client.create_collection("my_col2", embedding_function=None)

n = client.count_collections()
print(f"集合数量:{n}")
# 集合数量:2


说说我的看法
全部评论()
没有评论
关于
本网站专注于 Java、数据库(MySQL、Oracle)、Linux、软件架构及大数据等多领域技术知识分享。涵盖丰富的原创与精选技术文章,助力技术传播与交流。无论是技术新手渴望入门,还是资深开发者寻求进阶,这里都能为您提供深度见解与实用经验,让复杂编码变得轻松易懂,携手共赴技术提升新高度。如有侵权,请来信告知:hxstrive@outlook.com
其他应用
公众号