{"canonical":"https://www.readysetcloud.io/blog/allen.helton/agent-memory-needs-more-than-vector-search/","categories":["ai"],"contentText":"Over the past couple of years I’ve tried several times to teach an AI agent how to impersonate my voice. While I haven’t been super successful, it’s been a wonderful learning opportunity for me to dig into RAG, agent memory, and other advanced topics.\nI’ve iterated quite a bit on my original design, which was mostly just stuffing entire blog posts into the system prompt as examples. I built an interesting version of agent memory that had its shortcomings. Then I swapped over to Oracle and designed a hybrid search with three layers of memory, which was objectively better. But it still wasn’t perfect.\nActually, the memory part is pretty good. It’s the retrieval part that was leaving something on the table. It’s taken me a long time to figure out where I was going wrong with loading relevant few-shot examples. Let’s see if you can spot the problem:\nSELECT id, content, topic, VECTOR_DISTANCE(embedding, :q, COSINE) AS distance FROM posts WHERE user_id = :userId AND platform = :platform AND is_deleted = 0 ORDER BY distance FETCH APPROX FIRST :k ROWS ONLY Did you find it?\nNo? Well, that’s because there’s nothing technically wrong with it. It finds the memories most similar to the question. The problem is that the most similar memory isn’t always the right memory.\nMemories pile up. After a few years of writing, my store has a post where I argued for one approach and a newer one where I’d changed my mind. It has posts on event-driven architecture, event sourcing, and event streaming that are all very similar to each other in the embedding space. It has a post on serverless pricing and one on cold starts that use the same vocabulary but make completely different points. As far as similarity scores go, these groups are basically the same thing. It just comes down to a date, a specific term, or the actual point of the post.\nSo retrieval worked, but my context was garbage.\nThe standard fix is a reranker, which helps to improve relevance instead of just similarity.\nWhen you embed a memory, the model has no idea what will be asked of it later. Your query gets embedded and a similarity score is calculated against the embeddings in the store. A cross-encoder (a specific type of reranking model) looks at the query and memory together and scores them by how well a memory answers the question. This means the relevance is higher, and it often leads to better results.\nBut there’s a catch (of course).\nA cross-encoder runs at query time, once per candidate, on every request. That means added latency to your query. 🙃\nBut if the relevance of the results is meaningfully better, that makes up for the increased latency, right? To a point, yes. If the latency is so high the user experience suffers, it’s probably not worth the cost. But if it’s relatively minor, then absolutely.\nSo I built a benchmark to evaluate the tradeoff. I started from the evaluation plan in Oracle’s production RAG evaluation guide, which compares keyword, vector, and hybrid retrieval against judged queries. Then I pushed it one step further and put a reranker on top of each, to see how much relevance I got for the latency cost. For the reranker, I used the BGE reranker loaded into the database as an ONNX model, so scoring happens directly in the query with no extra service to call. 🔥\nLet’s dive in. The results are fascinating.\nBefore I continue, thank you to Oracle for sponsoring this post. Opinions are my own.\nFilter first, then rank RAG is cool because it involves multiple kinds of lookups. You have lexical retrieval, which is great for identifiers, codes, and names. You also have vector retrieval, which finds semantically similar data via embeddings. I benchmarked both on their own and together as a hybrid search, fusing the two ranked lists with reciprocal rank fusion (RRF).\nBoth paths run the same scope filter:\nWHERE TENANT = :tenant AND (OWNER_ID IS NULL OR OWNER_ID = :owner) AND (EXPIRES_AT IS NULL OR EXPIRES_AT \u003e SYSDATE) I share this because the WHERE clause handles eligibility, which is different from relevance. I don’t want an expired memory coming back in my results, like a 2025 take on building AI agents that I’ve since moved past. And I definitely don’t want another tenant’s data showing up (security nightmare). Oracle’s guide calls these filters mandatory for multi-tenant data. They run before ranking, so everything downstream only deals with memories that are allowed to be there.\nSetting the baseline Before I could tell whether a reranker was worth anything, I needed to know the baseline values for these queries. In addition to latency, I also needed to know the normalized discounted cumulative gain (nDCG) and recall. Capturing these metrics lets us objectively measure relevance alongside latency.\nThe baseline measurements were taken with K=10, meaning I took the top 10 results.\nRetrieval nDCG@10 Recall@10 Median latency Vector 0.753 0.865 20.0 ms Lexical 0.727 0.740 4.7 ms Hybrid RRF 0.832 0.885 20.5 ms nDCG measures how well a ranking puts relevant items near the top, with higher being better. It’s not a measure of how many answers were right, though. Vector’s 0.753 means that the rankings earned 75% of the score a perfect ordering would have, on average.\nRecall ignores order. It tells you whether the correct memories made it into the top 10 at all. A 0.865 means that about 87% of the memories that should have come back did. Both of these metrics are evaluated against a known ground-truth set of data in order to be objective. Meaning I know ahead of time what the results should have been given the corpus and the query I was benchmarking.\nYou can see from these results that lexical is fast, but it’s not super accurate. Vector catches more at the cost of latency, but still ranks them worse than I’d want. Fusing the two rankings had the best results, with minimal additional latency.\nFrom the baseline results alone, it looks like a hybrid search will be the retrieval method to beat if latency doesn’t drop off a cliff.\nThe enlightening reranker results Now to run the same benchmark with the reranker. In Oracle, it’s one more step at the end of the retrieval query. Retrieval narrows everything down to the top candidates, and PREDICTION() runs the cross-encoder against each one:\nSELECT id, PREDICTION(BGE_RERANKER USING :query || '\u003c/s\u003e\u003c/s\u003e ' || title || '. ' || content AS DATA) AS score FROM candidates ORDER BY score DESC FETCH FIRST 10 ROWS ONLY; The \u003c/s\u003e\u003c/s\u003e is the delimiter the model expects between a query and a passage. I’m only pulling back the ids and the reranked scores for the benchmark. And you might notice that the reranking is done entirely in the database query - super cool.\nConfiguration nDCG@10 Recall@10 Median latency Hybrid, no reranker 0.832 0.885 20.5 ms Hybrid + reranker, 10 candidates 0.832 0.885 585 ms Vector + reranker, 20 candidates 0.830 0.896 1,258 ms Surprisingly, adding the reranker to hybrid search didn’t do anything to nDCG at all. The best reranked score came from vector retrieval with 20 candidates, and it still scored lower than plain hybrid with no reranker, at 61 times the latency. It found slightly more of the right memories, but ranked them a little worse.\nI also tried giving the reranker more to work with. For vector retrieval, latency grew linearly with candidate count (1,258ms at 20 candidates and 2,604ms at 40), and nothing got meaningfully better.\nThat doesn’t mean the reranker did nothing, though. I dug a little deeper and took the change for each query and resampled them 2,000 times to see how much the average varies depending on which queries you run. 95% of the resampled averages landed between -0.079 and +0.073.\nAnd if we use that range, +0.073 would take hybrid from 0.832 to 0.905, which is a meaningful improvement. But conversely, the other end of the spectrum would drop it to 0.753, which is where plain vector search started.\nGiven the range, the results are inconclusive, at least with my sample size of 16 queries. The real effect could be a solid improvement, nothing at all, or even a loss, and 16 queries can’t tell those apart.\nThis ended up being a big realization for me. Relevance is measurable, but how confident you can be depends on how many judged queries you have. A couple thousand iterations of 16 queries is still only 16 queries, and that’s why the range is so wide.\nMeanwhile, the latency hit is pretty obvious. It was nearly 30 times slower for every request. Now I had to ask myself when that latency is worth it.\nWhere rerankers make sense Stay with me. Rerankers have lots of value despite what we’ve discussed so far. The underwhelming results are because I improved retrieval quality in a cheaper way. Let’s look again at the baseline compared to the reranked version:\nFirst stage No reranker 10 candidates 20 candidates 40 candidates Lexical 0.727 0.800 0.801 0.809 Vector 0.753 0.821 0.830 0.814 Hybrid RRF 0.832 0.832 0.829 0.817 Intuitively, the weaker retrieval methods got the most out of reranking. Lexical and vector each gained about 0.075, while hybrid gained nothing.\nWhich makes sense because a reranker doesn’t retrieve anything. It reorders (or re-ranks 😄) what it’s given. If your first stage is already pulling the correct memories into the pool in a sensible order, there’s nothing left to fix.\nReranking vector retrieval with 20 candidates left me a little curious because it made the top five less alike. I measured it a couple of different ways (with embeddings and with plain word overlap) and verified the numbers. Thinking about it though, that is desired behavior. Remember that memory lookalike issue I talked about? The cross-encoder is separating the right memory from its lookalikes and pushing the lookalikes down. Which is exactly what I said I wanted.\nThe takeaway I want you to leave with is that you should rerank when you have one retrieval signal, and fuse when you have two. Doing both leaves you paying for additional latency you don’t need.\nMy new approach I started this wanting to know whether a reranker was worth the extra inference. After running through several rounds of benchmarks and tweaking queries, I’m not sure that was the right question. My approach to the problem is much more informed and now walks a different path than I expected.\nWhen building a RAG pipeline for memory, I’ll always filter first. Owner, expiration, whatever makes a memory eligible for you. Apply these in the query before anything gets ranked. Reducing the number of candidates means fewer memories to rank (plus, you know, security and quality and all that 😜).\nNext, I’d add a second retrieval signal before adding a reranking model. If you have vector search, add lexical. If you have lexical, add vector. Fuse the rankings with RRF. Doing that with my tests took nDCG@10 from 0.75 with vector retrieval to 0.83 and only slowed my query down by half a millisecond.\nAfter adding the second retrieval signal, I’d measure everything. Write down 15 to 20 real requests you’d give your agent, making sure to include some difficult ones and edge cases. Things like asking for exact keyword lookups, or for advice on something you’ve changed your mind on a couple of times. Keep those handy. The guide I was following says to rerun them every time you change chunking, the embedding model, or the reranker, and it’s right.\nThen you can try a reranker against your baseline. Try a few different candidate counts (:n in the query above). Check whether the quality difference is consistently higher across your queries and decide if the additional latency is worth the quality difference.\nLuckily for me, what I built in my original post doesn’t need to change much. I still have the same tables and the same reflection loop. I’m really just updating the database query that decides which memories the agent sees.\nTry it yourself The benchmark and instructions to try it all yourself are available in GitHub. It spins up Oracle AI Database Free in Docker, generates a synthetic corpus of 480 memories and 16 judged queries, loads the ONNX models, and runs the whole thing.\nEverything we’ve discussed today ran on Free in Docker, capped at two CPUs and 2 GB. It’s an intentionally small box, because we aren’t running a capacity benchmark and none of these numbers should be used for sizing anything.\nIf you want to run this kind of evaluation on your own retrieval, start with Oracle’s production RAG evaluation guide. It covers more than what we talked about here, like grading the generated answer and testing freshness, tenant isolation, and when the agent should refuse to answer. It’s pretty solid.\nAgent memory needs more than vector search. Turns out, just using vector search means you’re asking a single similarity score to decide what’s relevant, current, and yours. Fortunately for all of us, “more” is a lot cheaper than I assumed.\nHappy coding!\n","date":"2026-09-29T00:00:00Z","description":"Vector search is great, but it only goes so far. I benchmarked several ways to improve relevance with surprising results.","image":"https://assets.readysetcloud.io/agent_memory_needs_more_than_vector_search_feature.jpg","inLanguage":"en-US","lastmod":"2026-09-29T00:00:00Z","readingMinutes":11,"section":"blog","tags":["sponsored","oracle","rag"],"title":"Agent memory needs more than vector search","url":"https://www.readysetcloud.io/blog/allen.helton/agent-memory-needs-more-than-vector-search/","wordCount":2194}