Zzong's Notes

Home

❯

Retrieval

❯

indexing

❯

Approximate Nearest Neighbor

Approximate Nearest Neighbor

2026년 6월 14일1 min read

Approximate Nearest Neighbor

We can split ANN algorithms into three distinct categories; trees, hashes, and graphs. HNSW slots into the graph category.

B) Github Repos

  • N2
  • GitHub - milvus-io/milvus: Vector database for scalable similarity search and AI applications.
  • GitHub - criteo/autofaiss: Automatically create Faiss knn indices with the most optimal similarity search parameters.

링크된 언급

3
faiss

Faiss 페북 (현 메타) 에서 만든 ANN 라이브러리. ColBERT 논문에서 빠른 retrieval 을 위해 faiss IVFPQ 버전을 사용했다고 한다. 네이버에서는 4ms 도 느리다고 판단하고, 보다 빠른 검색을 위해 Hnswlib 을 사용하는 것 같다. B) Faiss의 핵심 동작 원리: IVF, PQ Indexing IVF (Inverted File Index): 검색 공간을 줄...

HNSW

...le Small World (HNSW) graphs are among the top-performing indexes for vector similarity search (Approximate Nearest Neighbor). HNSW는 현재 가장 널리 쓰이는 ANN 알고리즘 중 하나로, 그래프 기반 알고리즘입니다.

scann

(1) Scann 구글에서 만든 ANN 라이브러리. 2022 년 기준 벤치마크 상으로 가장 좋은 성능을 내고 있다. B) (2) 튜닝하기 데이터가 100k 개 이상일 경우, AH 로 점수를 계산하고 rescore 절차를 거쳐야 한다. AH 로 점수를 계산할 때, dimensionsperblocₖ 은 2 로 설정하자. 파티셔닝 시에 numleaves 는 데이터포인트 개수의 제곱근 수 (squa...

함께 보면 좋은 글

HNSW

HNSW Hierarchical Navigable Small World (HNSW) graphs are among the top-performing indexes for vector similarity search (Approximate Nearest Neighbor).

scann

(1) Scann 구글에서 만든 ANN 라이브러리. 2022 년 기준 벤치마크 상으로 가장 좋은 성능을 내고 있다. B) (2) 튜닝하기 데이터가 100k 개 이상일 경우, AH 로 점수를 계산하고 rescore 절차를 거쳐야 한다.

faiss

Faiss 페북 (현 메타) 에서 만든 ANN 라이브러리. ColBERT 논문에서 빠른 retrieval 을 위해 faiss IVFPQ 버전을 사용했다고 한다. 네이버에서는 4ms 도 느리다고 판단하고, 보다 빠른 검색을 위해 Hnswlib 을 사용하는 것 같다.

ANN

ANN ANN(Approximate Nearest Neighbor)은 정확한 nearest neighbor를 전수 계산하지 않고, 충분히 가까운 후보를 빠르게 찾는 검색 방식이다.

milvus

Milvus Milvus는 대규모 vector search를 위한 vector database다.

Hnswlib

Hnswlib HNSW 기반 C++ 라이브러리 B) References.

Approximate Nearest Neighbors Oh Yeah

What is Annoy spotify 에서 만든 라이브러리로 유사한 벡터들을 빠르게 찾아주는 ANN 라이브러리.

Non-Metric Space Library

Non-Metric Space Library Non-Metric Space Library (NMSLIB) is an efficient cross-platform similarity search library and a toolkit for evaluation of similarity search methods.

locality sensitive hashing

Locality Sensitive Hashing References.

IVF

IVF IVF(Inverted File Index)는 vector space를 여러 centroid 또는 cluster로 나눈 뒤, query와 가까운 cluster 안에서만 후보를 찾는 ANN indexing 방식이다.

  • Approximate Nearest Neighbor
  • B) Github Repos