Old model
source embedding space
EmbedFlow
migration layer
New model
target embedding space

Upgrade embedding modelswithout rebuilding your index.

EmbedFlow lets a new embedding model serve traffic immediately while your target index is built progressively in the background.

02 — WHY

A model upgrade shouldn't require rebuilding a billion vectors first.

But today, it does.

Re-embedding a large corpus can take hours, days, or weeks before the new model can fully replace the old one.

Why wait for the backfill?

Upgrade first. Backfill as you go.

Billion-document values are projections from measured Qwen3-8B H100 throughput of 106.98 documents/sec.

TRADITIONAL MIGRATION10 BILLION DOCUMENTS
UPFRONT COST BEFORE SERVING
25,960H100 HOURS
1,080SINGLE-GPU DAYS
~$85KENCODING COMPUTE
~82 TBRAW FP16 VECTORS

Full backfill required before migration completes.

EMBEDFLOWNEW MODEL
SERVING
FIRST.
deploy

Backfill progressively after the new model is serving.

03 — THE BEHAVIOR

Use the new model.
before the index is ready.

We determine when your old index can safely support your new model.

OLD WORLDOld model stays live.

The existing index keeps serving while the next model is prepared.

OLD MODELEXISTING INDEX100% populated
NEW MODELTARGET INDEX0% ready
EMBEDFLOWmigration layer
NEW MODELWAITING
TARGET INDEX0% READYPREPARING
EVIDENCE, NOT INTERNALS

Backfill happens after deployment, not before it.

63migrations studiedsource → target transitions
+19.8msmeasured warm p50 overheadmeasured while serving traffic
<1msmigration machinerysmall overhead at serving time

For teams that are ready to move

Make the upgrade
without the pause.

Request early access