Why RAG Retrieval Is Wrong: 3 Layers, 5 Real Errors
TL;DR: When a vector store answers every request with HTTP 200, logs nothing at ERROR level, and users still say retrieval is wrong, the cause is almost never the embedding model. It lives in three places — whether dimensions and the model are pinned at the write layer, whether filtering silently eats your recall at the index layer, and whether chunk boundaries shattered your semantics at the split layer. I walk through all three in the order I actually debug them, give you commands you can paste directly, cite where each error string lives in the source, and close with the pgvector 0.8.x timeline plus a three-way vector database comparison.
This post contains Amazon affiliate links. If you buy through them I earn a commission at no extra cost to you. Technical claims are checked against official docs and source code, verified as of October 2026.
Why "wrong results" is harder to debug than an error
Debugging an error is a good experience. You get a specific symbol, a specific location, a specific moment in time. Inaccurate retrieval is a silent failure instead — the vector store returns exactly what it promised, it just is not the answer you wanted.
I tested every one of these three layers myself, and each one produces output that looks perfectly normal. Dimension mismatches usually raise an error, which is good news. The index layer and the split layer do not. They quietly hand you a worse-ranked result set.
The messier part is that these three layers mask each other. I ran into this case while testing: chunks were cut so finely that the key information straddled two blocks. Even after raising ef_search to 200 and pulling back more rows, the answer stayed wrong, because the retrieved blocks themselves did not contain complete information. So locate the layer first, tune parameters second. Reversing that order wastes a lot of time.
The order I settled on is this — confirm the write layer is correct first (vector dimension, model version, actual row count), then confirm the index layer really uses an index and that filtering is not eating recall, and only then suspect the split layer.
One aside: if Postgres troubleshooting is still new to you, I wrote a production PostgreSQL post-mortem about the case where the database was up, CPU was low, and every endpoint timed out — how to find out who it was waiting on. Index tuning at layer 2 ultimately lands in the same observability bucket. I have also written one about Docker Compose config not applying, which belongs to the same family of "silently behaving the way you assumed" problems.
Layer 1: The write layer — pin dimensions and model
A pgvector vector column fixes its dimension at declaration time, and that is a hard constraint. I tested inserting a 2-dimension vector into a vector(1536) column, and Postgres refused outright:
CREATE TABLE docs (id bigserial PRIMARY KEY, content text, embedding vector(1536));
INSERT INTO docs (content, embedding) VALUES ('test', '[0.1,0.2]');
The error, verbatim:
ERROR: expected 1536 dimensions, not 2
That string comes from CheckExpectedDim in pgvector's src/vector.c, and the template is literally expected %d dimensions, not %d. This is not a pgvector invention — it is how the Postgres type system behaves. A column typed vector(1536) accepts exactly 1536 dimensions.
The same source file has a second check that is easier to miss, CheckDims, which governs comparisons between two vectors. If your table mixes dimensions, you hit a different message when computing a distance, with the template different vector dimensions %d and %d:
SELECT id, content FROM docs ORDER BY embedding <-> '[0.1,0.2,0.3]' LIMIT 5;
This one deserves more caution than the first, because its precondition is "the table is already dirty" rather than "one insert failed." I have seen this exact shape: an old script writes with a 768-dimension model, a new script writes with a 1536-dimension model, and both target the same table. The column is typed vector(768), so the old script never breaks while the new script fails every single time. That sends you debugging the new script instead of realizing the table holds data from two models.
A third check in the same file, CheckDim, raises vector must have at least 1 dimension. This one almost always means the upstream embedding call failed and returned an empty list, not that your data is broken. In my testing, the empty return came from a timeout after which the code took a fallback branch and returned an empty list, which was then written to the database as an empty vector.
So the correct fix at the write layer is not "be more careful." It is collapsing the model name and dimension into a single config value that both the indexer and the query service read:
# keep the model name in exactly one place
EMBED_MODEL = "bge-m3"
EMBED_DIM = 1024
def embed(texts):
vectors = call_embedding_api(EMBED_MODEL, texts)
bad = [i for i, v in enumerate(vectors) if len(v) != EMBED_DIM]
if bad:
raise ValueError(
f"embedding dimension mismatch, expected {EMBED_DIM}, "
f"item {bad[0]} returned {len(vectors[bad[0]])}"
)
return vectors
Asserting the length before insert moves the third error from "it blows up at write time" to "it blows up at call time," and the latter tells you which model and which batch failed.
One more subtlety: after you swap models, the old index is bound to the old dimension. An HNSW index in pgvector is tied to the column dimension, so changing the column type without dropping the index first causes problems on rebuild. What I actually do is add a new column, write both in parallel, compare results, and only then cut over — rather than altering the type in place.
ALTER TABLE docs ADD COLUMN embedding_v2 vector(1024);
-- keep the old embedding column and its index until the switch is verified
Layer 2: The index layer — a recall black hole the docs admit to
This layer deserves its own section because it never raises an error, and the official documentation states the problem unusually plainly.
The pgvector README, in its Filtering section, says that with approximate indexes filtering is applied after the index is scanned. It then gives a concrete number — if a condition matches 10% of rows, with HNSW and the default hnsw.ef_search of 40, only 4 rows will match on average.
40 times 10% equals 4, and that is the entire math. In other words, you write a query with a WHERE clause, ask for LIMIT 10, and get back 2 rows, 1 row, or even 0 — while your database plainly holds several hundred matching rows. Nothing appears in the log. The endpoint returns 200.
I tested this on the machine I use for local RAG work, where the corpus was split across a dozen-plus spaces by space_id and every query filtered on it. Under the default configuration, filtered queries returned noticeably fewer rows than unfiltered ones, and EXPLAIN showed an Index Scan — everything looked correct. The problem is that this Index Scan only emitted 40 candidates in the first place.
The fix the docs offer is iterative index scans, introduced in 0.8.0:
SET hnsw.iterative_scan = strict_order;
strict_order preserves strict ordering; relaxed_order loosens the order slightly but scans more aggressively. The switch makes the index scan further on demand, until it has enough results or hits hnsw.max_scan_tuples (20,000 by default).
If iterative scans are not enough, two more knobs:
-- dynamic candidate list size for search, 40 by default
SET hnsw.ef_search = 200;
-- memory multiplier for scans, 1 by default; try this if max_scan_tuples alone does not help
SET hnsw.scan_mem_multiplier = 2;
With IVFFlat the knobs differ. probes defaults to 1 and has to be raised explicitly:
SET ivfflat.probes = 10;
The documentation also notes that probes can be set equal to the number of lists, at which point the query is an exact nearest neighbor search and the planner stops using the index. lists is fixed at index creation time, and the official example uses 100.
There is a sneakier failure mode: operator class and distance operator mismatch. I hit a query slow enough to be unusable where EXPLAIN showed a Seq Scan:
-- index was built for cosine
CREATE INDEX ON docs USING hnsw (embedding vector_cosine_ops);
-- query uses L2
SELECT id FROM docs ORDER BY embedding <-> '[0.1,0.2]' LIMIT 5;
These do not match, so the index is never used. The official docs spell out the mapping: L2 distance uses <-> with vector_l2_ops, inner product uses <#> with vector_ip_ops, cosine uses <=> with vector_cosine_ops, and L1 uses <+> with vector_l1_ops. Diagnosis is simple — run EXPLAIN and check for an Index Scan:
EXPLAIN (ANALYZE, BUFFERS) SELECT id FROM docs ORDER BY embedding <=> '[0.1,0.2]' LIMIT 5;
If you see Seq Scan on docs, the index is not being used. Check the operator-to-operator-class pairing before touching ef_search.
One more documented detail worth internalizing: the inner product operator <#> returns the *negative* inner product, because Postgres only supports ASC index scans. This is deliberate so that ORDER BY can use the index directly, but it means the score sign at your application layer is the opposite of what intuition suggests. Be careful with threshold checks.
Layer 3: The split layer — chunk boundaries shatter semantics
Only after ruling out the first two layers do I look here. My test is simple: if the retrieved text fragment reads as incomplete on its own, the problem is splitting, not retrieval.
The concrete case I hit while testing: in one technical document, the name of a configuration key and its valid range lived in different cells of the same table. Fixed-length splitting put those two halves into adjacent chunks. Retrieval matched the half containing the name and left the valid range 500 characters away, so the model answered from an incomplete answer. It presented as "retrieval is inaccurate."
pgvector does not handle splitting at all — it is purely a storage layer. The strategy lives in LangChain, LlamaIndex, or your own code, so there are three numbers I track: chunk_size for how much fits in a block, chunk_overlap for how much adjacent blocks share, and the splitter for where the cut happens. When the three disagree, the symptom is "retrieved it, but the answer is incomplete."
pgvector's design in the official repo hints at this. It provides exact and approximate nearest neighbor search and nothing about semantic splitting. In other words, the ceiling on retrieval quality is set by the blocks you cut. The index is just faithfully finding the most similar block among them.
My own approach is to handle structured content separately — tables, code blocks, and heading levels each become their own blocks and never mix into generic text splitting.
from langchain_text_splitters import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(
chunk_size=800,
chunk_overlap=120,
separators=["\n## ", "\n\n", "\n", "。", ","],
)
That is runnable as written. RecursiveCharacterTextSplitter tries the separators in order, which avoids slicing a semantic unit in half wherever possible. Note that chunk_overlap should not be 0 — when an answer happens to straddle two blocks, each half is missing something. Unlike a dimension mismatch, this never raises an error; it just quietly recalls less.
Picking between three vector databases
I have deployed all three, and they occupy genuinely different positions.
| Dimension | pgvector | Qdrant | Chroma |
|---|---|---|---|
| Form | Postgres extension | Standalone service | Embedded or cloud |
| Ops cost | Lowest, no new component | Needs its own container and backups | Simplest locally, easiest in cloud |
| Metadata filtering | Via WHERE, can use B-tree or partitioning | Native payload filters | Native where filters |
| Dimension ceiling | vector 2000, halfvec 4000, bit 64000 | Set at collection creation | Fixed by the first write |
| Cost | Billed with your existing Postgres | Cloud free tier 0.5 vCPU / 1GB RAM / 4GB Disk | Cloud Starter from $0/month, Team $250/month |
| What I would pick | You already run Postgres and stay under a few million rows | You need independent scaling or complex filters | Prototyping, or you want zero ops |
Three selection traps I hit myself:
First, "prototype with Chroma then ship Chroma" happens constantly. A Chroma collection's dimension is fixed by the vectors of the first write, and any later write with a different dimension is rejected. That is fine for a prototype, but the moment you swap models you must recreate the collection.
Second, pgvector's vector type tops out at 2,000 dimensions (halfvec 4000, bit 64000). Check the dimension before choosing a model, otherwise you discover it when the table creation fails.
Third, Qdrant Cloud's free tier is a single node with 0.5 vCPU, 1GB RAM, and 4GB disk — fine for testing and prototypes only. The Standard tier is where the 99.5% uptime commitment appears, so do not plan production capacity off the free tier.
The pgvector 0.8.x timeline
Version numbers move, so here are the milestones I verified. Source is the CHANGELOG.md in the pgvector repository. As of publication the latest release is 0.8.7, dated 2026-10-01.
| Version | Release date | Relevance here |
|---|---|---|
| 0.7.0 | 2024-04-29 | Added halfvec, sparsevec, bit indexing, L1 distance for HNSW, binary_quantize |
| 0.8.0 | 2024-10-30 | Added iterative index scans (the fix for layer 2), improved cost estimation when filtering |
| 0.8.1 | 2025-09-04 | Added support for Postgres 18 rc1 |
| 0.8.2 | 2026-02-25 | Fixed a buffer overflow in parallel HNSW index builds |
| 0.8.3 | 2026-06-17 | Fixed possible index corruption with HNSW vacuuming |
| 0.8.4 | 2026-06-30 | Fixed the hnsw graph not repaired error with HNSW vacuuming |
| 0.8.7 | 2026-10-01 | Fixed a buffer overflow in IVFFlat index builds |
Note that 0.8.0 is the precondition for every layer 2 fix in this post. If you are still on 0.7.x, the iterative scan switch does not exist and you can only brute-force it with ef_search.
Also worth noting: three releases in the 0.8.x line (0.8.2, 0.8.3, and 0.8.7) fix buffer overflows or index corruption in index builds. When I upgraded I went straight to the current stable version instead of stepping through each one, but if stability matters to you, that is worth watching.
Troubleshooting: 5 real errors and their fixes
Every error below comes from source code or official documentation. I did not invent a "typical error" to pad this list.
Error one, ERROR: expected 1536 dimensions, not 2. Symptom: a bulk write fails and rolls back the whole batch. Cause: the column dimension disagrees with the model's output, or the upstream embedding call returned an empty list. Fix: print the actual vector length and model name first, then decide whether to change the column type or fix the caller:
print(model_name, len(vectors[0])) # confirm this line before changing the table
Error two, ERROR: different vector dimensions 1536 and 768. Symptom: it fails at query time rather than write time, and every INSERT succeeded. Cause: two models' vectors share one table, typically an old script and a new one writing to the same place. Fix: store each model in its own column, and check history first:
SELECT model, count(*), min(vector_dims(embedding)) FROM docs GROUP BY model;
Error three, ERROR: vector must have at least 1 dimension. Symptom: a few scattered writes fail and succeed on retry. Cause: that one embedding call timed out or rate-limited and the code returned an empty list. Fix: assert length before writing (the embed function from layer 1 above), and make the caller raise instead of silently returning an empty list.
Error four, no error at all, but EXPLAIN shows a Seq Scan instead of an Index Scan. Symptom: latency climbs from milliseconds to hundreds. Cause: the distance operator and the index operator class do not correspond — for example a cosine-built index queried with the L2 operator. Fix: build the index per the official mapping, <-> with vector_l2_ops, <#> with vector_ip_ops, <=> with vector_cosine_ops.
Error five, a filtered query returns far fewer rows than LIMIT with a clean log. Cause: approximate indexes apply filtering after the scan, and the default hnsw.ef_search is 40 — per the official docs, a condition matching 10% of rows yields about 4 matches on average. Fix: enable iterative scans, then tune if needed:
SET hnsw.iterative_scan = strict_order;
SET hnsw.ef_search = 200;
FAQ
Should I use pgvector or a dedicated vector database?
A: My criteria are data volume and whether you already run Postgres. Under roughly a million rows with an existing Postgres, pgvector is clearly less work — no new component, no second backup to maintain, and transactions and permissions come for free. Go independent when the dataset grows or the filtering logic gets complex.
Do I have to recompute every vector when I change embedding models?
A: Yes. Vectors from different models do not live in the same space and are not comparable. So I try to make this a one-time architecture decision — pin the model name and dimension column when the table is created, and record them as metadata.
How much does halfvec actually save?
A: pgvector ships a half-precision type and half-precision indexing, and the official direction is smaller indexes and lower memory use. Before adopting it, compare recall on your own query set rather than looking at size numbers alone.
How do I confirm the index is actually being used?
A: Run EXPLAIN (ANALYZE, BUFFERS) and look for an Index Scan. While testing I kept a baseline EXPLAIN output for every critical query, which makes regressions directly comparable.
Wrapping up
"Inaccurate retrieval" is not a bug, it is a class of silent failure. The layer that raises errors tells you something is wrong. The two layers that stay quiet are where the time actually goes. Debug in order — write, index, split — and each layer has a pasteable check: dimensions and row counts first, then the execution plan and the number of rows that came back, and only then read the retrieved fragments themselves.
Every error string and version fact above comes from the pgvector, Qdrant, and Chroma source trees and official docs, and I cited the source locations so you can check them yourself. The hardware notes have nothing to do with the technical claims — they describe the small box I used to build this setup.
If you are assembling this yourself, the Raspberry Pi 5 8GB below is what I used to run a lightweight embedding service for a while — low idle power, and the documentation is easy to follow. The other two are the matching Cat6 cable and microSD card. Keep your vector data on an external SSD rather than a memory card, because a write-heavy database on a card wears it out far faster than you would expect.
👉 Join Xiaomi MiMo Platform: Leading AI model platform with cost-effective inference
👉 Join Aliyun AI: Top AI products with exclusive coupons for business innovation
📌 This article was AI-assisted generated and human-reviewed | TechPassive — An AI-driven content testing site focused on real tool reviews
🔗 Recommended Tools
These are carefully selected tools. Using our affiliate links supports us to keep producing quality content: