Why we rebuilt instead of porting
The old system was a two-node LangGraph controller with LangChain for model calls, in-process BM25 for retrieval with embeddings stored in pgvector on Cloud SQL, and LangSmith for traces. The app ran on one small VM behind Cloudflare. We had three reasons to rebuild it rather than port it.
One ecosystem. The VM and the database were Google Cloud, but orchestration came from LangGraph, tracing from LangSmith with its own API key, and the edge from Cloudflare. We wanted the whole system on Google Cloud: Agent Platform for retrieval and sessions, ADK for the agent, Cloud Run and Firebase for serving, Cloud Trace for observability, and one service account authenticating all of it.
Full-size embeddings. The data had been updated and re-crawled, so every vector had
to be re-embedded regardless. gemini-embedding-001 produces 3,072 dimensions, but
pgvector's HNSW index tops out at 2,000 dimensions for full-precision vectors, so the old build
truncated them to 1,536. With the store being rebuilt anyway, we wanted one that keeps all 3,072.
A new platform. In April 2026 Google relaunched Vertex AI as Gemini Enterprise Agent Platform, with managed retrieval and agent sessions. We wanted to run it on a real workload rather than a demo.
The old backend was archived, and a six-step spec barred the new build from importing it; only the React frontend carried over. Each step had an acceptance check, and the crawler, index, agent and gateway were built between 11 and 13 September.
Choosing a retrieval backend on Agent Platform
We compared five Google options for putting documents behind the agent, as of September 2026:
| Option | Why it did or did not fit |
|---|---|
| Vector Search 1.0 | Bills node-hours for the VM hosting the index. Google's cheapest worked example is $68 a month, more than the whole old stack. |
| RAG Engine | Spanner mode keeps 100 processing units running and never scales to zero. Serverless mode was Preview, us-central1 only, with no CMEK. |
| Agent Search data store | Managed chunking, embedding, hybrid search and reranking, with no fixed cost. It cannot run four arms fused by rank or take our own sparse vectors, and chunking is fixed at creation. |
| Gemini File Search | Cheapest: free storage and query embeddings, a few dollars to index this corpus. It is not on Agent Platform; it needs the Gemini Developer API and an API key, and this project uses none. |
| Agent Retrieval (chosen) | One batchSearch carries every sub-search with its own filter and fuses them by RRF. Exhaustive KNN needs no index and no capacity unit. Sparse vectors are a first-class query type. |
Agent Search came closest, but it runs one hybrid search per call. Our design runs four arms fused by rank, and two of them handle specific lookups, where the question names an exact entity. On the old system, the golden set's lookup category scored 0.933 with them and 0.190 without.
Agent Retrieval runs that design as it is, and it had a fallback. If its text search did not rank like BM25, we could store BM25 weights as sparse vectors, whose dot product with the query is the BM25 score.
What replaced what
| Job | LangGraph build | Agent Platform build |
|---|---|---|
| Orchestration | LangGraph: a two-node controller loop, up to three retrieval passes, with map-reduce over sub-questions in plain Python | ADK 2.9.0: router, rewriter and answer agents in sequence, plus at most one deterministic second round inside the search tool |
| Model calls | LangChain ChatGoogleGenerativeAI on Vertex, global endpoint | ADK calling Gemini directly, same endpoint |
| Vector store | pgvector on Cloud SQL as an embedding store, no vector index | An Agent Retrieval collection, exhaustive KNN |
| Embeddings | gemini-embedding-001 truncated to 1,536 and normalised by hand | gemini-embedding-001 at 3,072, generated by the collection on import |
| Lexical retrieval | BM25 in process (rank_bm25, then bm25s) | The collection's full-text search, terms joined by OR |
| Fusion | RRF in Python, k=60, α=0.48 | RRF inside batchSearch, equal weights |
| Reranking | Ranking API, semantic-ranker-default-004, with a hedged second request | Ranking API, same model |
| Conversation state | Our own session module, 30-minute expiry | Agent Sessions |
| Serving | FastAPI on a GCE e2-medium, Cloudflare in front | FastAPI on Cloud Run behind Firebase Hosting, scaling to zero |
| Tracing | LangSmith | Cloud Trace, from ADK's own OpenTelemetry spans |
| Secrets | A database connection string and a LangSmith API key | None |
The models did not change. Both builds route and answer with gemini-3.5-flash-lite and
rewrite queries with gemini-3.7-flash.
Browser (React 19, Vite 7)
| Firebase Anonymous Auth + App Check
v
Firebase Hosting ---------------------- static files + SPA
| rewrite /api/**
v
Cloud Run gateway verifies the ID token, streams SSE
| 2 instances x 4 concurrent turns
v
ADK agent, in-process router -> rewriter -> answer
|
+--------------------+---------------------+
v v v
Agent Retrieval Ranking API Agent Runtime
batchSearch + RRF top 20 Sessions
66,060 objects (no deployed code)
^
| importDataObjects
Cloud Storage <- crawler: 16,111 pages, two completeness gates
The agent runs in-process inside the Cloud Run gateway, the same process that verifies the token and streams the answer. The Agent Runtime instance holds no deployed code; it exists only to hold Sessions. Cloud Run bills only while it serves requests and Hosting stays inside its free tier, so the floor of the bill is collection storage.
Four arms in one request
The rewriter splits each question into one to four sub-queries, 1.9 on average when it searches.
Each sub-query becomes one
batchSearch call carrying up to four searches, fused inside the service and billed as
one request.
question
|
v
rewriter up to 4 sub-queries
|
v per sub-query, concurrently, ONE batchSearch:
| dense semanticSearch top 200 always
| lexical textSearch, terms joined by OR top 200 always
| lookup semanticSearch + field $eq only if the question names one
| lookup semanticSearch + field $eq (a second field, same rule)
| combine RRF, equal weights top 200
v
Ranking API top 20 per sub-query
|
v
admission at most 2 chunks per page, at most 8 pages,
then each page expanded to one contiguous span
|
v
answer model reads the merged passages, cites every URL it used
The filter makes a lookup exact, not the search type. Each lookup arm is
an exact $eq filter on one field, so only chunks carrying that value can come back.
batchSearch accepts only vector, semantic and text searches, so the filter has to ride on
one of them. The spec paired it with textSearch, and it scored 0.000. The collection's
text search is conjunctive: a chunk matches only if it holds every term, and a question phrased as a
sentence nearly always carries a word, like "what" or "is", that the target chunk lacks. A plain
question naming a record matched 0 of 66,060 objects. semanticSearch
takes the same filter and never comes back empty because of wording; switching to it took lookup
recall to 1.000 on 30 test entities, with ground truth from the corpus's own fields.
Reweighting the arms does nothing. A lookup arm is only as big as its filter. The largest single identifier holds 163 chunks, so none of the 8,481 identifiers can fill a 200-slot arm, and only 3 of the 1,250 values in the other lookup field can. On three probes, everything a lookup arm returned was already in the fused list, and weights from 1.5x to 5x changed list membership by exactly zero. The reranker then scores all 200 fused results, its maximum, so RRF's ordering never reaches the answer.
Reranking is a separate call. The batchSearch API has an optional Vertex reranker that runs after RRF, but it accepts only the
semantic-ranker-fast model, and sent without RRF the service rejects it with
Ranker is required. Fusion happens in the batch, and the Ranking API runs after it.
The reranker scores relevance and ignores redundancy, so for one broad question its first four hits were four consecutive chunks of the same page, and keeping the top six left only three pages to cite. Admission now walks the reranker's top 20, keeping at most two chunks per page and at most eight pages, then expands each page into one contiguous span. Across six representative queries, distinct pages rose from 31 to 46. The spans are wide on purpose: on one sub-query, 12 admitted chunks across 8 pages became 28 delivered passages, because a page that argues across several chunks only reads correctly with the middle present.
The lexical arm, three times
Dense retrieval misses exact strings. To measure how much without a golden set, we lifted a nine-word span from each of 100 random chunks and checked whether that chunk's page came back in the top 200. Dense found 0.47, BM25 found 0.65, and 0.22 were found by BM25 alone. Verbatim spans flatter a lexical matcher, but the gap still made a lexical arm mandatory. Getting one that worked took three attempts.
| Attempt | How it works | What happened |
|---|---|---|
1. textSearch as specified | The collection's full-text search, default dialect | Conjunctive. Scored 0.000 on every real question while passing keyword probes. |
| 2. BM25 as sparse vectors | Each object stores its BM25 term weights with corpus-wide IDF. The query holds 1.0 per known term, so the dot product is the BM25 score. | Exact: 0.413190 against bm25s's 0.41319 on a test fixture. Also a full scan, at 4.7 s for any query. |
3. textSearch, structured OR | One subquery per term the BM25 vocabulary knows, joined by OR | 0.70 s at the median. Finds 63 of 100 known items where BM25 finds 65. |
The sparse build did what it promised: 69,394 terms, about 125 non-zero weights per object, scores matching local BM25. Then production traces showed a warm turn taking 14.4 s at the median, 7.8 s of it in the search tool. The arms run in parallel, so the slowest one sets the floor, and the slowest by far was the sparse arm.
filter query, no search 270 ms the service floor
dense, 8 terms, top_k 200, all fields 1,631 ms
sparse, 1 term, top_k 10, url only 4,710 ms
sparse, 8 terms, top_k 10, url only 5,894 ms
sparse, 8 terms, top_k 200, all fields 5,049 ms
A one-term query costing as much as an eight-term query, with top_k making no
difference, is a full scan: all 66,060 sparse vectors are touched on every query. An inverted index
reads only the postings for the query's own terms, which is why an in-process BM25 index answers the same query in 0.2 ms.
The fix was the text search we had already rejected. It answers from an index, and a structured query
can join terms with OR instead of AND. That query is Preview and exists only in the
v1beta surface, so the serving image pins google-cloud-vectorsearch at 0.11.3; the v1beta types in 0.11.2
lack it. We ran every golden case under both arms, back to back:
| Sparse vectors (before) | Text search, OR (now) | |
|---|---|---|
| Pass / fact recall | 59/59, 1.000 | 59/59, 1.000 |
| Search tool, avg / p95 | 8.52 s / 15.74 s | 3.21 s / 5.61 s |
| Turn, p50 / p95 | 13.77 s / 32.40 s | 10.38 s / 22.16 s |
| Tokens and ranking per turn | $0.0103 | $0.0102 |
The search tool was faster on all 56 turns that searched, by 4.8 s at the median. The gap grows with the number of sub-queries, because parallel scans contended with each other: at four sub-queries, 15.7 s against 5.6 s at the median. The dense arm, at 1.6 to 1.9 s depending on the run, is now the floor. The sparse vectors are still in the collection, one environment variable away.
Moving BM25 back in process would have been as fast end to end, but its hits would still need their fields fetched from the collection, and the container would carry a second index to rebuild with every corpus.
Letting the collection do the embedding
The old build chunked, embedded, normalised and inserted every vector itself. Agent Retrieval can generate the embedding on import from a text template, so the new build hands it text and this template:
{title} | {category} | {identifier} | {content}
The template carries page identity. In the old corpus the identity prefix was added before splitting, so only each page's first chunk had it, and 89% of chunks did not know what page they were on. The old build compensated with three separate patches, at embedding, BM25 and reranking. The template does it once: every chunk is embedded with its page's title, category and identifier. Chunks run about 2,000 characters with 200 of overlap, roughly 500 tokens.
Vectors are stored at the full 3,072 dimensions. Exhaustive KNN has no dimension ceiling, and the full size adds $0.08 a month in storage.
Building and querying the collection held most of the undocumented behaviour. Each of these failed outright, which at least made them quick to find:
- Agent Retrieval is its own API,
vectorsearch.googleapis.com, with its own IAM roles. Theaiplatformpath returns 404, androles/aiplatform.usergrants nothing there. - The import format is
{id, data, vectors}, and object ids cap at 63 characters. Of our 66,060, 13,519 were longer, so long ids keep their first 51 characters and their last 12 (a six-character hash of the URL plus the chunk ordinal), verified collision-free. - The data schema forces
additionalProperties: false: every field must be declared, and a null is not a string. - A filter map holds exactly one key. Two conditions side by side return 400; wrap them in
$and. - A vector field added after creation is invisible to
batchSearch.UpdateCollectionaccepts it,getCollectionlists it andsearchDataObjectsuses it, butbatchSearchanswersSearch field ... not found in collection schema. That is the one call a multi-arm design depends on, and it cost us a full collection rebuild.
The import's long-running operation reports no progress, its success count stays at zero after
success, and op.result() raises on an import that worked. Only a quota metric in Cloud
Monitoring told a slow import from a dead one. When an expired credential killed our poller midway,
rerunning the importer would have started a second import and re-embedded all 66,060 objects at full
price, so it now attaches to any import already in flight.
Check a new project's quotas before anything else. Ours got 50,000 embedding tokens a minute, against
1,000,000 on a project with history, which stretches an import that took under an hour to a projected 11 hours. A Cloud
Quotas request for 1,000,000 came back with '50000' was granted.
Three agents in a row instead of a loop
The spec translated the LangGraph design into ADK almost node for node: a router, a schema-constrained
rewriter, a ParallelAgent for retrieval, a LoopAgent as the coverage gate,
then map, answer and reduce agents. We built three.
SequentialAgent
router gemini-3.5-flash-lite in scope / direct / refuse
rewriter gemini-3.7-flash <= 4 sub-queries
answer gemini-3.5-flash-lite calls search_corpus, answers, cites
The loop and map-reduce solved problems the new retrieval no longer has. The old build needed a coverage gate because its first round could come back thin, and map-reduce because four sub-questions sharing one rerank budget dropped about 4 of 48 facts per run, a different four each time. Now the reranker cuts each sub-query to 20, admission builds a context of about 20,000 tokens, and the answer model reads it in one pass. We added nodes only when a measurement called for one, and the one that did asked for something smaller than a loop.
ParallelAgent could not have done the fan-out anyway: it takes a fixed list of
sub-agents, and the rewriter decides the number of sub-queries at runtime. They run concurrently
inside the tool instead.
Two ADK behaviours cost us time:
- State reaches a prompt only through
{placeholders}. The answer agent's instruction referred to a field of the rewriter's output in prose, so the model never saw its value. ADK substitutes only{key}placeholders that name a state entry, here{route}and{rewrite}; any other mention of state stays plain text, with no error. SequentialAgentis deprecated in 2.9.0 in favour ofWorkflow, which cannot yet be used as anLlmAgentsub-agent. It still works.
Gemini 3.x models are served only from the global endpoint. us-west1,
us-central1 and us-east4 all answered NOT_FOUND, and
GET publishers/google/models/{id} returns 404 for every id, including the ones that work.
For most of a day we believed the models did not exist. Only a real generateContent call
against global settles it. The collection stays in us-west1; the two
locations are unrelated.
Model tiers do not carry across generations. gemini-2.5-flash-lite passed three of four
smoke cases, then failed an in-scope question, twice answering without retrieving at all.
gemini-3.5-flash-lite passed all four and took the answer node. The rewriter was
specified on gemini-3.8-flash, but on 18 September that model answered 8 of our 37
requests in 25 minutes and refused the rest with 429. ADK does not retry, so each refusal was a failed
turn. gemini-3.7-flash answered every request and is the rewriter now.
A second hop without a loop
Multi-hop looked healthy until we tested hops whose two halves share no retrieval neighbourhood: an identifier whose linked entity also appears under a different group. Those scored 0 of 6. The one earlier case with this shape had passed on luck: both halves sat in the same group, and dense retrieval pulled them in together.
The cause was not a missing loop. The lookup arms fire only on an entity the question names, and a second hop by definition does not name it:
"what else is that one listed on" -> [dense, lexical]
"what else is <name> listed on" -> [dense, lexical, lookup:<name>]
The lookup arm worked the whole time and never got to run. So the tool now runs one more round for the two entities the first round found most often and the question did not name, as a plain filter on that field with no ranking. It is skipped only when the first round turns up no such entity. One round, not a loop: nothing we have seen needs a third.
before after
cross-group second hop 0/6 6/6
latency 7.0 s 7.3 s
passages 42 55
The old LangGraph controller made this call with a model, judging whether the evidence was enough and naming what was missing. The new rule is code: more predictable, less general. It covers every multi-hop case we have, and a question that truly needs a third hop will need something more.
Two smaller defects surfaced on the way:
- Enumeration is not ranking. "What else is this one listed on" went through rerank-then-keep-the-top, which dropped whole groups: a filter on the name returns all eight identifiers one entity is listed under, across three groups, and the ranked path kept three from one group. Enumeration now uses the filter directly, which is exact and skips the reranker: 14.6 s down to 7.8 s.
- Picking follow-up entities alphabetically is picking at random. One identifier's first round returned five different entities, and
sorted()[:2]chose two of the wrong ones. Follow-ups are now weighted by their share of the first round.
One known limit remains. Enumeration runs only on the second hop, so a question that names an entity directly stays on the ranked path, capped at eight pages per sub-query. An entity listed under 30 identifiers gets 4 of them that way, and many more when reached indirectly. No golden case names an entity listed that widely, so the score does not show it.
Sessions and identity without a login
The old build kept conversations in about 60 lines of its own session code, keyed on a UUID the server issued and then accepted back from the request body unchecked, so anyone holding another user's id could use it. The new build keeps conversations in Agent Sessions, scoped by Firebase Anonymous Auth. The gateway takes the scope from the verified ID token, never from the request body. The identity is per browser and per device, and with no login, a "Forget me" control deletes the current session and then the anonymous identity; older sessions under that identity become unreachable rather than deleted.
Two wiring traps. Plain google-adk cannot reach Agent Sessions without the
google-adk[gcp] extra. And ADK's SSE mode sends every partial, then a final event with
the whole answer again, so a client that appends deltas shows the answer twice unless the gateway
drops the last one. Both events also carry the call's token usage, which doubled the count shown
under each answer. Prices copied from an older model were three to six times too low, so even doubled, the
displayed cost ran about 40% under the real one for a typical turn.
Anonymous sign-up is an open endpoint, so identities are free to mint in bulk and a per-user limit counts nothing. Three controls each bound a different part of the cost:
| Control | What it bounds |
|---|---|
| Firebase App Check (reCAPTCHA Enterprise) | Whether the caller is this app in a real browser, not who it claims to be |
| 2,500-character message cap, rejected with 422 before any model runs | The cost of one request |
| Cloud Run at 2 instances × 4 concurrent turns | How fast money can be spent, not how much |
Before App Check, a valid ID token took one unauthenticated sign-up call and about three seconds of
curl. Now the gateway answers 403. Enforcement went in two steps so that browsers holding
a cached bundle from before App Check would not break: the gateway logged every verdict without
rejecting anything, and enforcement went on once the logs showed real browsers arriving with
ok.
The service runs with --allow-unauthenticated because a Hosting rewrite calls Cloud Run
without IAM credentials, so the gateway's own checks are the whole boundary: App Check first, then the ID
token. There are no
secrets: no API key, no connection string, only Application Default Credentials through one service
account.
Tracing without LangSmith
ADK instruments itself with OpenTelemetry, so a turn already produces about 14 spans
(invocation, invoke_agent, call_llm,
generate_content, and execute_tool under the agent that called it). They
carry the prompt as the model received it, the response, the search tool's arguments and returned
passages, and per-call token counts. All we added was an exporter, about a hundred lines with the two fixes below.
Two fixes made it one readable trace per turn. Cloud Run opens its own span from
X-Cloud-Trace-Context, and without a propagator that reads it, ADK's spans start a
separate trace: you can see that a turn took 36 seconds but not what it spent them on. Reading that header and opening a server span under it joins the two. The ASGI instrumentation that opens
it also opens a span per streamed chunk, 12 in one measured turn and hundreds in a long answer, so its receive and send
sub-spans are excluded.
Two limits apply. The exporter truncates an attribute at 65,536 bytes, and one retrieval response measured 72,657 characters, so the largest payloads arrive clipped. And spans hold the user's question and the passages retrieved for it, a second copy of user text that Cloud Trace keeps for 30 days.
Performance and cost, before and after
| LangGraph build | Agent Platform build | |
|---|---|---|
| Infrastructure cost per month | $52.46: about $24 for the VM and $28 for Cloud SQL | About $0.35; the floor is 1.14 GiB of collection storage |
| Tokens and ranking per turn | $0.0088, across 234 graded turns | $0.0102, across 59 golden cases |
| Corpus | 11,129 documents, 100,515 chunks | 16,111 pages, 66,060 chunks |
| Embedding dimensions | 1,536 | 3,072 |
| Always-on infrastructure | A VM and a database | None |
| Secrets | 2 | 0 |
A production turn costs more than tokens and ranking. Over seven live turns through the deployed gateway, read from its own token counts:
| Line | Per turn |
|---|---|
| Gemini tokens | $0.0058 (median of six ordinary turns; range $0.0039 to $0.0081) |
| Ranking API | $0.0020 |
| Agent Retrieval reads | $0.0002 |
| Agent Sessions events | $0.0015 |
| Cloud Run | $0.0007 |
| Total | about $0.010 |
Model tokens are the only line that moves much. The seventh turn, a question that enumerates
everything one entity is listed on across groups, reached 86,714 input tokens and $0.032, three times a
typical turn, so the mix of questions sets the bill more than the average does. At about $0.010 a turn, 10, 100 and 1,000 turns a day
come to about $3, $31 and $307 a month. On 1 January 2027,
gemini-3.7-flash's introductory $0.75 and $3.75 per million tokens double; at about 450
tokens in and 130 out, a turn rises by about $0.0008.
The first ceiling is a quota, not Cloud Run. Embedding allows this project 5 requests a minute, every sub-query spends one, and a turn averages 1.8, so sustained traffic near three turns a minute hits 429, and two four-query turns in the same minute already do. Raising it takes a Cloud Quotas request, and on a new project the answer may be no.
Latency got worse, though no single pair of numbers compares it cleanly. The old build answered a one-search turn in 6.9 s at the median and a four-search turn in 11.9 s. The new one answers the golden set at 10.4 s p50 and 22.2 s p95 from a workstation, plus about 1.5 s of session writes in production. In 13 September production traces, with the earlier rewriter model, the rewriter took 2.3 s and each of the other three model calls about a second; the search tool, since reworked, now averages 3.2 s. The first question of a session also pays Cloud Run's cold start, which the server cannot see: one request measured 17.3 s inside the handler while the user waited 34.0 s. The browser now times the round trip itself and shows that number.
Quality does not compare across the two builds. The re-crawl kept only the latest version of each record page, which moved essentially every record URL the old golden sets pointed at, so the new system has its own set: 59 cases in eight categories, every expectation mined from the corpus and checked against it before scoring. It passes 59 of 59 with fact recall 1.0.