The Insanity of Relying on Vector Embeddings: Why RAG Fails

Paper · Source
Retrieval-Augmented Generation (RAG)LLM Failure Modes

Wrong Tool for the Job

RAG fails in production because vector embeddings are the wrong choice for determining percentage of sameness. This is easily demonstrated. Consider the following three words:

King

Queen

Ruler

King and ruler can refer to the same person (and are thus considered synonyms). But king and queen are distinctly different people. From the perspective of percentage of sameness, king/ruler should have a high score and king/queen should be literally zero.

In other words, if the query is asking something about a “king” then chunks discussing a “queen” would be irrelevant; but chunks discussing a “ruler” might be relevant. Yet, vector embeddings consider “queen” to be more relevant to a search on “king” than “ruler.” Here are the vector similarity scores for queen and ruler when compared to king using OpenAI’s ADA-002 embeddings:

King

Queen: 92%

Ruler: 83%

Reranking

Reranking is perhaps the most recommended Advanced RAG strategy. However, as the RankRAG study shows, even using a fine-tuned model for reranking only results in a 54.2 score on EN.QA. Using general reranking models had an even worse score.

GraphRAG and Knowledge Graphs

A recent study on KG-RAG (RAG enhanced with Knowledge Graphs) showed an F1 score of 25% and an accuracy of 32% for CWQ dataset. Interestingly, Knowledge Graph RAG had a lower accuracy than regular embedding RAG (which had a 46% accuracy).

Real-World vs Hype

There simply is no study showing that vector embeddings, combined with dozens of Advanced RAG techniques, results in a reliable chatbot in production environments containing numerous documents. Moreover, the added latency of many Advanced RAG techniques makes them impractical for real-world chatbots—irrespective of the accuracy issue.

But even if larger models could overcome the problem, they would be slower and more expensive. In other words, they’d be too slow and too expensive for any practical purpose. Would companies pay more for a chatbot than for a person, when the chatbot would require up to a minute for each unreliable answer?

That’s the actual state of RAG. That’s the actual outcome of relying on vector embeddings.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Do structural constraints outperform deep architectures in recommendation systems? Why do embedding systems fail to capture task-relevant relationships? How should items be represented and indexed in recommenders? What causes retrieval-augmented generation systems to fail despite access to external knowledge? Do knowledge graphs offer advantages over embeddings for multi-hop retrieval? When do semantic similarity approaches miss structural retrieval failures? How should retrieval systems handle complex multi-step reasoning?