Filtered Vector Search: Pre-Filtering, Post-Filtering and HNSW Metadata Filters That Do Not Break Recall
Every demo of vector search runs on an unfiltered index, and every production system runs on a filtered one. The moment a retrieval-augmented generation (RAG) application has to answer “only from this customer’s documents, published after March, that this user may read,” the neat nearest-neighbour story falls apart. The query returns three results instead of ten, or ten results that are not the true nearest ten, or nothing at all while thousands of matching rows sit in the table. None of these failures raise an error. The system just quietly gets worse.
Filtered vector search is the problem of combining approximate nearest neighbour (ANN) retrieval with a structured predicate, and it is hard because graph indexes such as HNSW were built without any knowledge of your metadata. Their navigability is a property of the whole dataset, and a filter removes most of it at query time.
This article explains why filters break ANN indexes, does the over-fetch arithmetic, compares pre-filtering, post-filtering and in-traversal filtering, covers ACORN and how Qdrant, Weaviate, Elasticsearch and pgvector respond to the problem, and ends with a measurement script and a decision matrix.
What this covers: the connectivity failure mode, selectivity and correlation, the three filtering strategies, ACORN, vendor implementations, partitioning and multi-tenancy, a selectivity-driven planner, pgvector SQL, a recall-measurement harness, and a decision matrix.
Context and Background
Approximate nearest neighbour indexes trade exactness for speed. The dominant family for in-memory and memory-mapped search is the proximity graph, with Hierarchical Navigable Small World (HNSW) as the reference design: a layered graph in which a query greedily walks from a coarse top layer toward the neighbourhood of the query vector, then beam-searches the dense bottom layer with a candidate list of size ef_search. The internals, including how graph degree M and ef_construction interact with quantization, are covered in our guide to HNSW, DiskANN and IVF-PQ index internals. This article assumes that background and asks a narrower question: what happens when only a subset of the vectors is allowed to be returned.
The question matters because almost no real workload is unfiltered. Multi-tenant SaaS needs tenant isolation. Enterprise RAG needs access-control lists. E-commerce needs price, stock and category constraints. Document systems filter by date, language and source. In each case the filter is not an optimisation hint; it is a correctness requirement, and the result set must contain only rows that satisfy it.
Historically, the ecosystem treated this as a bolt-on. The original ANN libraries, and the first generation of vector databases built on them, exposed search(vector, k) and left filtering to the caller. Two workarounds emerged: filter the corpus first and search the survivors, or search first and filter the results. Both appear in the vendor documentation to this day. Elastic’s documentation, for instance, distinguishes a filter placed inside the knn clause, which is applied during the vector search so that k matching documents are returned, from a filter placed beside a knn query in a bool query, which is applied afterwards and can return fewer than k hits or none.
The research community caught up with the engineering reality in 2024. The ACORN paper by Liana Patel, Peter Kraft, Carlos Guestrin and Matei Zaharia (arXiv 2403.04871) framed hybrid search over vectors and structured data as a graph-traversal problem and reported 2 to 1,000 times higher throughput than prior methods at fixed recall. Since then the idea has spread: Lucene adapted an ACORN-1 style exploration, Weaviate made ACORN its default filter strategy from v1.34, Qdrant added an optional ACORN search mode from v1.16, and Elasticsearch enables an acorn heuristic by default for indices created in 9.1 or later, per the vendor documentation fetched for this article.
For teams running vectors inside Postgres, the story is related but different. pgvector 0.8.0 introduced iterative index scans precisely because filtered queries were returning too few rows. Our pgvector versus Qdrant ADR on hybrid search and reranking and the broader pgvector versus a dedicated vector database comparison discuss when each platform is the right home; this article is about what happens inside the index once you have chosen.
The thesis, which the rest of the article defends, is that filtered search is not one problem but three, distinguished by selectivity (how many rows pass), correlation (whether passing rows cluster near the query), and cardinality of the filter key (a few tenants versus millions of tags). A system that handles one regime well can fail silently in another, so the engineering task is to detect the regime per query and route accordingly.
Why Filters Break Approximate Nearest Neighbour Search
A filter breaks HNSW because the graph was wired by vector similarity alone. When a predicate removes most nodes, greedy traversal either runs out of passing neighbours and stalls, or must wade through huge numbers of failing nodes to find a few passing ones. The three strategies differ in where they apply the predicate: before the search, after it, or during traversal.

Figure 1: Three places to apply a metadata filter relative to the HNSW traversal, and the characteristic failure of each.
Figure 1 shows the three options as pipelines. Post-filtering runs the ANN search unchanged and discards failing hits. Pre-filtering resolves the predicate to a set of allowed identifiers (or scans the matching rows directly) and then searches only those. In-traversal filtering lets the graph walk continue through the full structure but admits only passing nodes to the result set. The first fails by returning too few results, the second by being slow or by stranding the walk, and the third by needing a graph that stays connected under the predicate.
Post-filtering and the over-fetch arithmetic
Post-filtering is the default behaviour of any index that does not know about your predicate. You ask for the top k by distance, then apply WHERE to what came back. If the filter passes a fraction s of rows (the selectivity) and passing rows are spread uniformly through the space, each returned candidate survives with probability s. To end up with k survivors you must retrieve roughly k / s candidates.
The numbers get uncomfortable quickly. With k = 10: at s = 50% you need about 20 candidates; at s = 10% about 100; at s = 1% about 1,000; at s = 0.1% about 10,000. HNSW search cost grows roughly with the candidate list size ef_search, so a 1,000-candidate beam is far slower than the default, and latency budgets are usually the first thing to break. These figures are expected values under a uniform-survival assumption, which is the best case. Real survivors are random, so a beam sized to the expected count still falls short of k roughly half the time; production over-fetch factors need a safety margin of two to three times.
Concretely, pgvector’s documented default hnsw.ef_search is 40. At 10% selectivity a query with LIMIT 10 receives 40 candidates of which about 4 survive. The query returns four rows. At 1% selectivity the expected survivors are 0.4, and the query frequently returns nothing. This is the “empty results” failure that motivated iterative scans in pgvector 0.8.0.
Why graph connectivity collapses under a selective filter
Over-fetching is the visible problem. The deeper one is topological. At the base layer an HNSW node typically holds on the order of 2M outgoing links (the hnswlib convention, which pgvector follows with its m parameter defaulting to 16, so about 32 links per node). Suppose a filter keeps a fraction s of nodes and, for the moment, pretend they are scattered uniformly. A traversal that only follows passing neighbours sees about 32 s usable links at each node.
A back-of-envelope random-graph model makes the cliff visible. The probability that a given node has no passing neighbour at all is (1 - s)^32. At s = 10% that is about 3%; at s = 5% about 19%; at s = 1% about 72%. Percolation theory says a random graph fragments into isolated islands when its average degree falls toward one, which here happens near s = 1/32, about 3%. Below that, the passing nodes no longer form one connected component, so a walk that enters one island cannot reach another no matter how large ef_search is. This is an illustrative model, not a measurement, since real HNSW graphs are not random graphs, but it predicts the shape practitioners observe: recall holds up through moderate selectivity, then drops off a cliff.
The consequence is counter-intuitive. Raising ef_search helps post-filtering up to a point, but it does nothing for a traversal confined to a disconnected subgraph. The walk terminates when its candidate queue is exhausted, and with an island of 50 passing nodes the queue empties long before it has explored the rest of the filtered set. The result is silent recall collapse: the index returns something plausible, only not the true nearest neighbours among the passing rows.
Correlation: the factor selectivity hides
Selectivity is only half the picture. Correlation between the filter and the query vector determines whether the passing rows are near the query or far from it. If a tenant’s documents cluster in their own region of embedding space and the query comes from that tenant, the passing rows are right there, the filtered neighbourhood is dense, and even aggressive filters behave well. This is the positively correlated case.
The negatively correlated case is the dangerous one. Imagine searching for products similar to a winter coat while filtering to category “swimwear.” The nearest vectors to the query are all coats, all excluded, and the passing rows sit in a distant region. The graph walk spends its whole budget on failing neighbours near the query and never reaches the swimwear cluster. Weaviate’s documentation describes its ACORN strategy as most useful exactly when the filter has low correlation with the query vector and is restrictive, and the ACORN paper’s evaluation includes workloads where predicates are uncorrelated with the embedding.
This matters for how you test. Benchmarks that draw filters and queries independently measure the uncorrelated case. A real multi-tenant workload is strongly positively correlated within a tenant and rarely hits the worst case, while a faceted e-commerce search with many independent filters can hit it constantly. Know which regime your traffic lives in before choosing a strategy, and measure filtered recall on your own query log, as the script later in this article does.
Pre-filtering: exact, but bounded by the matching set
Pre-filtering resolves the predicate first. With an inverted or B-tree index you obtain the set of matching row identifiers, then either compute distances to every match (brute force over the matching set) or search the ANN structure restricted to that set.
The brute-force variant is exact and, for small matching sets, fast. Scoring 5,000 vectors of dimension 768 is about 3.8 million multiply-adds, comfortably under a few milliseconds on a modern core, and entirely independent of graph connectivity. It degrades linearly: scoring a million matches costs a thousand times more. Qdrant’s own guidance, retrieved for this article, says pre-filtering gives exact results but becomes expensive when many points match. The sweet spot is therefore the narrow end of the selectivity range, which is precisely where HNSW is weakest, and that complementarity is the basis of the planner described later.
Weaviate documents the same trade: when a filter is very restrictive, about 15 percent of the dataset or less per its documentation, it switches to brute-force search over the matching subset, which is more efficient than HNSW traversal. The threshold is a per-collection setting rather than a universal constant, because it depends on vector dimension, quantization and hardware.
In-traversal filtering: the graph walks everything, returns only matches
The third approach changes what the traversal does at each node. The search still uses the full graph for navigation, so failing nodes act as stepping stones, but only passing nodes are added to the result heap. The walk terminates when it has found k passing results (or exhausted its budget), not when it has examined ef_search nodes of any kind.
This guarantees k results when enough matches exist and the walk is allowed to continue, which is why Elastic’s documentation says a filter inside the knn clause is applied “to ensure that k matching documents are returned.” The price is cost: a very selective filter means the walk examines many failing nodes per passing one, and in the worst case degrades toward a full scan. Weaviate calls this approach sweeping, and Qdrant’s filterable HNSW and the ACORN family are refinements aimed at making it both cheaper and more robust to disconnection.
How Production Systems Keep Filtered Recall Intact
The practical designs fall into four families: make the graph filter-aware at build time, make traversal predicate-aware at query time, keep scanning until enough rows pass, or avoid the problem by partitioning. Production systems usually combine two of them with a planner that picks per query.

Figure 2: ACORN-style predicate subgraph traversal. At each visited node the search expands to two-hop neighbours and keeps only those that satisfy the predicate.
ACORN: traversing the predicate subgraph
ACORN, from Patel, Kraft, Guestrin and Zaharia, starts from a clean observation. The ideal hybrid search would build an HNSW index over only the passing vectors, but the predicate is unknown until query time. ACORN approximates that ideal by traversing the “predicate subgraph,” the subgraph of the existing index induced by passing nodes, and by making that subgraph dense enough to stay navigable.
The paper describes two variants, per its text. ACORN-gamma builds a denser graph by gathering M x gamma candidate neighbours per node instead of M, where gamma is chosen as the inverse of the minimum selectivity the system should serve before falling back to pre-filtering. A predicate-agnostic pruning step keeps the index compact by dropping a candidate when it is already reachable through the two-hop neighbourhood of kept neighbours. At query time the search routes greedily as in HNSW but considers only neighbours that pass the predicate. According to the authors, the index is at most 1.3 times the size of an HNSW index and takes at most 11 times as long to build.
ACORN-1 is the cheap variant. It builds a plain HNSW graph and does the expansion at query time: at each visited node it collects one-hop and two-hop neighbours, applies the predicate, and truncates to M. The paper reports up to 5 times lower throughput than ACORN-gamma at fixed recall, but 9 to 53 times lower time to index, with an index at most 1.25 times HNSW’s size.
The headline results, again as reported by the authors, are 2 to 10 times higher queries per second than prior methods on earlier benchmarks, over 30 times on their new benchmarks, and over 1,000 times at 25 million vectors, all at 0.9 recall. These are the authors’ numbers on their own datasets; treat them as evidence of the mechanism rather than a prediction for your workload. The arXiv listing we checked did not name a peer-reviewed venue, although a Crossref record exists for the published version, so cite the preprint for the technical details.
The mechanism deserves restating, because it explains every vendor feature below. A two-hop expansion turns the 32 s usable-link count from the earlier model into roughly 32 x 32 x s candidate links, so connectivity survives to much lower selectivity. The cost is evaluating predicates on many more nodes per step, which is cheap if the predicate is a bitmap lookup and expensive if it is an arbitrary function.
Vendor implementations, as documented
Qdrant builds the filter into the index. Per its documentation, you create payload indexes on every field you filter on, ideally before ingesting data. Qdrant then adds extra edges between points that share a value of an indexed field, so a filtered walk does not get stuck; fields indexed after ingestion get these edges only after the HNSW index is rebuilt. Its query planner estimates, per segment, how many points match, and if the matches fit under full_scan_threshold (documented default 10,000 KB) it scores them directly for exact results; otherwise it searches the filterable graph. The documentation flags a risk zone where a small share matches but too many to score directly, notably with several strict filters. For that zone, v1.16 added an optional ACORN search-time mode, off by default and applied when the matching share is below max_selectivity (default 0.4). For multi-tenancy it recommends one collection with a tenant field marked is_tenant: true rather than a collection per customer.
Weaviate offers three filter strategies selected by a filterStrategy setting. Sweeping traverses from the entry point and skips failing nodes until enough results are found. ACORN, which has been the default since v1.34, ignores non-matching objects in distance calculations, evaluates multi-hop neighbourhoods and seeds extra matching entry points. A newer PathSeer strategy (v1.40 and later per the docs) is sweeping with ACORN-like neighbour checks as the result set fills; the documentation says it was fastest or near-fastest in most of Weaviate’s own benchmarks, with ACORN remaining faster for sparse, negatively correlated filters. Underneath all three sits the pre-filter allow list from the inverted index and the flat-search cutoff described earlier.
Elasticsearch and Lucene apply a filter inside the knn clause during the approximate search. Elastic’s engineering write-up describes a Lucene adaptation of ACORN-1: it explores only neighbours that satisfy the filter, extends to neighbours of neighbours when more than 10% of the immediate neighbourhood is filtered out, stops early once enough candidates are scored, and branches further with bounded effort for very restrictive or inversely correlated filters. The new path is used only when 40% or more of vectors are filtered out, since gains level off near 60%, and the article reports up to 5 times faster filtered kNN in Lucene 10.2. The reference documentation adds that for indices created in 9.1 or later the acorn heuristic is on by default and that very high recall targets may need a larger num_candidates.
Milvus is a case where we could not verify specifics from primary documentation during this run. It supports scalar filtering through bitsets and partition-based isolation, and the academic surveys on filtered ANN cover it, but consult the current Milvus documentation for exact index and iterator behaviour before relying on it.
pgvector takes the third route, scan until satisfied, and we cover it in detail next.
pgvector: iterative scans and the planner trap
In stock Postgres a vector query with a filter is two things: an index scan that returns approximate neighbours in distance order, and a WHERE clause applied to each row the scan returns. That is post-filtering by construction. Before 0.8.0 the HNSW scan produced at most hnsw.ef_search candidates, so a selective WHERE returned fewer rows than LIMIT asked for.
Version 0.8.0 added iterative index scans. With hnsw.iterative_scan set to strict_order or relaxed_order, the index continues scanning past the initial ef_search candidates until enough rows survive the filter or a limit is hit. According to the project documentation for v0.8.5, strict_order preserves exact distance order while relaxed_order allows slightly out-of-order results for better recall; the scan is bounded by hnsw.max_scan_tuples, which defaults to 20,000 and is approximate; and hnsw.scan_mem_multiplier (default 1, a multiple of work_mem) should be raised if increasing the tuple cap does not help. The same docs show a materialized CTE with ORDER BY distance + 0 on Postgres 17 or later to restore strict ordering from a relaxed scan, and advise placing any distance filter outside the CTE and other filters inside. IVFFlat has an equivalent ivfflat.iterative_scan and ivfflat.max_probes.
Two caveats follow. Iterative scanning reuses the unfiltered graph, so it is a bounded in-traversal strategy with no extra connectivity: a filter below the percolation cliff can hit max_scan_tuples and return fewer than LIMIT rows regardless. And Postgres, not pgvector, decides whether to use the vector index at all. For a highly selective predicate on an indexed column, the planner may correctly choose the B-tree and sort the survivors by distance, which is exact pre-filtering. For a moderately selective predicate it may choose the vector index and filter, which is where recall suffers. The planner knows nothing about the correlation between your filter and your query vector, so its row estimates can steer it wrong.
Partitioning and per-tenant indexes
When the filter key has modest cardinality and is always present, the cleanest fix is to stop filtering at all. Partition the data so that each tenant, region or product line owns a separate index, and a query touches only one. Because every index contains only passing rows, the connectivity problem disappears: a per-tenant HNSW graph is an ordinary unfiltered graph.
In Postgres this means declarative list partitioning with a vector index per partition, or partial indexes (CREATE INDEX ... WHERE tenant_id = 42) for a handful of large tenants. Partition pruning then routes the query. Elsewhere it means a collection or namespace per tenant, or Qdrant’s single collection with a tenant-aware index. The cost is operational: thousands of small indexes consume memory and file handles, each needs maintenance, and cross-tenant queries must fan out. Partitioning also does nothing for ad hoc filters such as price ranges and free-form tags, which brings us back to graph-level techniques. A common hybrid is to partition on the one high-cardinality isolation key and use filterable-graph features for everything else.
A selectivity-driven planner
All of the above converge on one architecture: estimate selectivity first, then pick the cheapest strategy that preserves recall. Qdrant and Weaviate both ship a version of it. If you are on a system that does not, you can build one at the application layer.

Figure 3: A selectivity-aware planner. Very selective filters go to exact scan, mid-range filters to a filter-aware graph search, and permissive filters to plain ANN with a post-filter safety net.
Figure 3 shows the decision flow. The planner obtains a cardinality estimate for the predicate, from a payload index count, a Postgres EXPLAIN row estimate, or a cached histogram. If the estimated matching set is small enough to brute-force within the latency budget, it scores the matches directly and returns exact results. If the filter passes most rows, a plain ANN search with a modest over-fetch is enough, because failing hits are rare. In between sits the danger zone where filter-aware traversal, ACORN or iterative scanning, earns its keep, with a bounded budget and a fallback to the exact path when the walk returns fewer than k rows.
Thresholds should be set empirically from the cost curve of exact scoring on your hardware and vector dimension, not copied. As reference points from the vendors, Weaviate’s flat-search cutoff documentation says about 15% of the dataset, Lucene engages its ACORN path at 40% filtered out, and Qdrant’s ACORN mode defaults to a max_selectivity of 0.4. They measure different things, so do not treat them as one number. The point is the shape: three regimes and two boundaries.
Hands-On: pgvector Filtered Queries and a Filtered-Recall Harness
The only way to know whether your filtered queries are healthy is to compare them with ground truth. This section gives a minimal schema, the pgvector settings that matter for filters, and a Python script that measures filtered recall against exact brute force. The SQL targets pgvector 0.8.0 or later; check SELECT extversion FROM pg_extension WHERE extname = 'vector'; first.
Schema, index and iterative scan settings
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE chunks (
id bigserial PRIMARY KEY,
tenant_id int NOT NULL,
lang text NOT NULL,
published date NOT NULL,
embedding vector(768) NOT NULL
);
-- Index the filter columns so the planner has a real pre-filter option
CREATE INDEX chunks_tenant_pub ON chunks (tenant_id, published);
CREATE INDEX chunks_hnsw ON chunks
USING hnsw (embedding vector_cosine_ops) WITH (m = 16, ef_construction = 64);
-- Per-session knobs
SET hnsw.ef_search = 100;
SET hnsw.iterative_scan = relaxed_order; -- or strict_order
SET hnsw.max_scan_tuples = 20000; -- documented default
-- Filtered query: planner may use the vector index (post-filter + iterative)
-- or the B-tree (exact pre-filter), depending on its row estimate.
EXPLAIN (ANALYZE, BUFFERS)
SELECT id, embedding <=> $1 AS dist
FROM chunks
WHERE tenant_id = 42 AND published >= DATE '2026-03-01'
ORDER BY embedding <=> $1
LIMIT 10;
Read the plan. If you see Index Scan using chunks_hnsw with a Filter: line and Rows Removed by Filter far larger than the returned rows, you are paying the over-fetch tax. If you see a bitmap or index scan on chunks_tenant_pub followed by a sort, Postgres chose exact pre-filtering. Both are correct; what you must ensure is that the choice matches the selectivity regime. For a handful of very large tenants, partial indexes remove the question:
CREATE INDEX chunks_hnsw_t42 ON chunks
USING hnsw (embedding vector_cosine_ops) WHERE tenant_id = 42;
A partial index is used only when the query’s WHERE clause implies the index predicate, so keep the literal tenant_id = 42 in the query text. To force exact results for a diagnostic run, disable the index scan for one transaction with SET LOCAL enable_indexscan = off;, which makes Postgres fall back to a sequential scan and sort.
Measuring filtered recall
Filtered recall at k is the fraction of the true top-k among passing rows that the approximate query returned. The ground truth is the exact nearest neighbours restricted to the filter, computed by brute force. Measure it over a representative sample of real queries and filter values, and report it by selectivity bucket, because an average hides the cliff.
import numpy as np, psycopg
from collections import defaultdict
K = 10
conn = psycopg.connect("dbname=rag")
def run(sql, params):
with conn.cursor() as cur:
cur.execute(sql, params)
return [r[0] for r in cur.fetchall()]
def selectivity(tenant, since):
with conn.cursor() as cur:
cur.execute("SELECT count(*) FILTER (WHERE tenant_id=%s AND published>=%s)::float"
" / count(*) FROM chunks", (tenant, since))
return cur.fetchone()[0]
APPROX = """SELECT id FROM chunks WHERE tenant_id=%s AND published>=%s
ORDER BY embedding <=> %s::vector LIMIT %s"""
def exact(tenant, since, qvec):
with conn.transaction(force_rollback=True), conn.cursor() as cur:
cur.execute("SET LOCAL enable_indexscan = off") # force exact path
cur.execute(APPROX, (tenant, since, qvec, K))
return [r[0] for r in cur.fetchall()]
def bucket(s):
for edge in (0.001, 0.01, 0.05, 0.2, 0.5):
if s < edge: return f"<{edge}"
return ">=0.5"
stats = defaultdict(lambda: {"hits": 0, "possible": 0, "short": 0, "n": 0})
for tenant, since, qvec in load_workload(): # your sampled query log
truth = set(exact(tenant, since, qvec))
if not truth: continue
got = run(APPROX, (tenant, since, qvec, K))
b = bucket(selectivity(tenant, since))
st = stats[b]
st["hits"] += len(truth & set(got))
st["possible"] += len(truth)
st["short"] += int(len(got) < min(K, len(truth)))
st["n"] += 1
for b, st in sorted(stats.items()):
print(f"{b:>8} recall@{K}={st['hits']/st['possible']:.3f} "
f"underfilled={st['short']/st['n']:.1%} queries={st['n']}")
The script reports two numbers per selectivity bucket. Recall tells you whether returned rows are the right ones; the underfill rate tells you how often the query returned fewer rows than it should have, the failure most dashboards miss. load_workload() is a placeholder you must supply, ideally from logged production queries so that correlation between filters and vectors is realistic. Run it before and after toggling hnsw.iterative_scan, changing ef_search, or adding partial indexes, and gate deploys on the lowest-selectivity bucket rather than the mean.
Two details keep the comparison honest. Use the same distance operator in the exact and approximate queries, and count a tie at the k-th distance as a hit, since tied vectors make the true top-k ambiguous. For very large tables, sample the workload rather than the data, and compute exact answers on a read replica so the sequential scans do not disturb production.
Decision matrix
| Situation | Selectivity | Best strategy | Why | Watch out for |
|---|---|---|---|---|
| Per-tenant RAG, tenants are modest in size | Very low (under ~1%) | Partition or per-tenant index, or exact scan | Matching set is small; no connectivity issue | Index sprawl, cross-tenant queries |
| Enterprise ACL filter, user sees a slice | 1% to 20% | Filter-aware graph (ACORN, filterable HNSW) or pgvector iterative scan | Walk needs bridges across failing nodes | Underfilled results; test on real ACLs |
| Faceted e-commerce with several filters | Varies, often negatively correlated | ACORN-style or sweeping with planner fallback | Filters and query vector disagree | Per-query cost spikes |
| Date or language filter that most rows pass | Over ~50% | Plain ANN plus post-filter and small over-fetch | Failing hits are rare | Over-fetching more than needed |
| Mixed, unpredictable workload | Anything | Cardinality-based planner with exact fallback | One strategy cannot cover all regimes | Stale statistics |
Use the matrix as a starting hypothesis, then confirm with the harness. If the platform choice is still open, our comparison of pgvector, Qdrant and LanceDB maps these capabilities against operational cost.
Trade-offs, Gotchas, and What Goes Wrong

Figure 4: The four recurring failure modes of filtered vector search and the signal that exposes each.
Figure 4 groups the failure modes that recur across engines. The first is the disconnected subgraph, discussed above: a filter selective enough to fragment the passing nodes makes recall collapse, and raising ef_search does not repair it. The signal is recall that is fine in aggregate but near zero for the rarest filter values.
The second is the underfilled result. Many clients treat “fewer rows than LIMIT” as normal, so nobody notices. In RAG the effect is subtle: the language model receives less context and answers from a thinner evidence base, or from the one irrelevant chunk that happened to pass. Log the returned count against the requested k and alert on the underfill rate.
The third is stale or misleading statistics. A planner that routes on cardinality estimates is only as good as the estimates. Postgres row estimates after bulk loads without ANALYZE can be badly off, and a predicate combining several columns multiplies independent selectivities that are not independent. Qdrant explicitly warns about the risk zone with multiple strict filters for the same reason. When the planner picks the vector index for a predicate that is far more selective than it believed, you get the over-fetch tax and an empty result.
The fourth is correlation blindness, covered earlier. Nothing in a selectivity number tells you whether the passing rows are near the query. Test with filters deliberately chosen to be far from the query distribution.
Several further gotchas deserve a mention. Filter-aware structures are built at index time: Qdrant’s extra links exist only if the payload index existed when the graph was built, so adding a payload index to a loaded collection requires a rebuild to take full effect. Deletions matter too: heavily soft-deleted graphs behave like filtered graphs, and Qdrant’s guidance suggests ACORN for collections with many deleted points. Quantization adds error before the filter even matters, so a filtered query on a heavily quantized index often needs more rescoring candidates; see the quantization discussion in our index internals guide.
There are anti-patterns to avoid. Over-fetching by a fixed large constant (say LIMIT 1000 then filter in the application) hides underfill for common filters and wastes work for rare ones. Setting max_scan_tuples to a very large value to “fix recall” converts a bounded query into an unbounded one and moves the problem to tail latency. And per-customer collections at tens of thousands of customers is an operational failure, which is why Qdrant documents the single-collection tenant pattern.
Finally, be honest about exactness. Pre-filtering with brute force is the only strategy here that is exact; everything else is approximate, and ACORN’s authors report throughput at 0.9 recall, not perfect recall. If your compliance story requires that every authorised document be findable, approximate filtered search alone is the wrong tool, and you need a lexical or exact fallback. Security filtering has a separate rule: a filter that enforces access control must be applied inside the data path, never only in the application after retrieval, since a post-filter can leak the existence of documents through result counts and timing.
Practical Recommendations
Start by measuring, because the right design depends on a distribution you probably have not looked at. Pull a week of production queries, compute each query’s filter selectivity, and plot the histogram. Most workloads turn out to be bimodal: many queries hit a tiny slice (one tenant, one product line) and a minority are broad. Those two groups want different strategies, and treating them with one index setting is the usual mistake.
For the narrow slice, prefer isolation: partitions, partial indexes, per-tenant namespaces or a tenant-aware index, or a plain exact scan when the matching set is small. For the broad group, a plain index with a modest over-fetch is fine. For the awkward middle, adopt whatever filter-aware traversal your engine provides, and bound its work. On pgvector, enable hnsw.iterative_scan, set max_scan_tuples deliberately, and run ANALYZE after bulk loads. On Qdrant, create payload indexes before ingestion and consider the ACORN option for combined strict filters. On Weaviate, keep the default ACORN strategy unless your own benchmark shows another strategy wins for your filters. On Elasticsearch, put filters inside the knn clause, not beside it.
Then close the loop. Gate releases on filtered recall in the lowest-selectivity bucket, alert on underfilled results, and re-run the harness after every embedding-model change, since a new model reshapes the embedding space and therefore correlation.
Checklist:
- Histogram of real filter selectivity from the query log.
- Filter columns indexed; payload indexes created before bulk load where the engine needs them.
- Filters inside the ANN call, never only post-hoc in application code.
- Iterative scan or filter-aware traversal enabled, with a bounded budget.
- Exact fallback when fewer than
krows return or the estimated match set is small. - Planner statistics refreshed after bulk loads; thresholds tuned on your hardware.
- Filtered recall and underfill rate reported per selectivity bucket in CI.
- Access-control filters enforced in the database, not the client.
Frequently Asked Questions
What is filtered vector search?
Filtered vector search returns the nearest neighbours of a query vector among only those records that satisfy a metadata predicate, such as a tenant, date range or category. It is hard because approximate indexes such as HNSW were built from vector similarity alone, so applying a predicate before, after or during traversal each carries a trade-off between recall, latency and the number of results returned.
Should I use pre-filtering or post-filtering for vector search?
Use pre-filtering, meaning exact scoring of the matching set, when the filter is very selective and the matching set is small. Use post-filtering with a small over-fetch when the filter passes most rows. For the middle range, use in-traversal filtering such as ACORN, Qdrant’s filterable HNSW or pgvector iterative scans, because naive post-filtering returns too few rows and naive pre-filtering gets slow.
Why does my vector search return fewer results than the limit?
Almost certainly post-filtering. The index returns its top ef_search candidates, the database discards the ones failing your WHERE clause, and fewer than LIMIT survive. In pgvector enable hnsw.iterative_scan (0.8.0 and later) so the scan continues, or raise ef_search. In other engines move the filter into the vector query clause so it is applied during the search.
What is ACORN in vector search?
ACORN is a method from Patel, Kraft, Guestrin and Zaharia (2024) for predicate-agnostic hybrid search on HNSW. It traverses the predicate subgraph, expanding to two-hop neighbours so the walk stays connected when many nodes fail the filter. Its authors report 2 to 1,000 times higher throughput than prior methods at fixed recall. Variants now appear in Weaviate, Qdrant, Lucene and Elasticsearch.
How do I measure recall for filtered queries?
Sample real queries and filter values, compute the exact top-k restricted to the filter by brute force, and compare it with the approximate result. Report recall and the underfill rate, the share of queries returning fewer than k rows, per selectivity bucket rather than as one average, since the failures concentrate at low selectivity. The harness in this article does this for pgvector.
Is it better to partition by tenant or filter by tenant?
Partitioning avoids the connectivity problem because each index contains only one tenant’s vectors, but it costs memory, maintenance and cross-tenant fan-out, and it suits a modest number of large tenants. Filtering within one index scales to many small tenants, provided the engine has a tenant-aware or filter-aware structure. Qdrant, for example, recommends a single collection with a tenant field over a collection per customer.
Further Reading
Internal:
- pgvector vs Qdrant: hybrid search, reranking and production RAG ADR
- HNSW vs DiskANN vs IVF-PQ: vector index internals and quantization
- pgvector vs a dedicated vector database
- pgvector vs Qdrant vs LanceDB
External:
- ACORN: Performant and Predicate-Agnostic Search Over Vector Embeddings and Structured Data (arXiv 2403.04871)
- Qdrant: A Complete Guide to Filtering in Vector Search
- pgvector iterative index scans documentation
- Elastic: Filter approximate kNN results
By Riju — about
