如果你要获取当前租户+数据库下面所有的向量集合信息,可以使用 list_collections 函数。
函数签名如下:
client.list_collections(limit=None, offset=None)其中:
limit 最多返回多少个集合。默认上限 100 条,超过 100 必须分页
offset 分页偏移量,从第 offset 条之后开始取
例如:
import chromadb
client = chromadb.EphemeralClient()
def create(name:str, num:int) -> None:
col = client.create_collection(name=name, embedding_function=None)
col.add(
ids=[f"qa{i}" for i in range(num)],
documents=[f"这是问题{i}" for i in range(num)],
metadatas=[{"tag":f"tag{i}"} for i in range(num)]
)
# 准备数据
create("my_col1", 2)
create("my_col2", 10)
create("my_col3", 5)
# 查询集合信息
cols = client.list_collections()
for c in cols:
print(f"集合名:{c.name}, 数量 {c.count()}")
# 输出:
# 集合名:my_col3, 数量 5
# 集合名:my_col2, 数量 10
# 集合名:my_col1, 数量 2注意,list_collections() 返回的是 Collection 对象列表,按创建时间从旧到新排序。
默认最多返回 100 个。超过要用分页:
# 前 100 个
batch1 = client.list_collections(limit=100, offset=0)
# 后 100 个
batch2 = client.list_collections(limit=100, offset=100)
# 从第 50 个开始取 20 个
subset = client.list_collections(limit=20, offset=50)如果你只想要数量,请使用 count_collections() 函数。 统计当前 database 内一共有多少个集合,直接返回整数,不需要遍历全部集合,性能远优于 len(client.list_collections())。例如:
import chromadb
client = chromadb.EphemeralClient()
# 准备数据
client.create_collection("my_col1", embedding_function=None)
client.create_collection("my_col2", embedding_function=None)
n = client.count_collections()
print(f"集合数量:{n}")
# 集合数量:2