Colourbox · Skyfish

ML anddata atColourbox

Building search that understands content — and running the models for it inside European jurisdiction.

Kasper Rømer GrøntvedAI Engineer
Structural draft. The flow is right; the specifics still need to come out of the source deck.
Before we start

Who is talking

Kasper Rømer Grøntved

Kasper Rømer Grøntved

AI Engineer at Colourbox · Odense

Where I come from

A PhD in multi-robot systems at SDU, with a stay at Carnegie Mellon's Robotics Institute — mostly about how one person stays in control of many autonomous things at once.

What I do now

Search at Colourbox and Skyfish: a stock library anyone can browse, and a DAM where every customer only ever sees their own material. Two very different search problems, one team.

And what I keep arguing about

European data sovereignty — running the models on our own infrastructure rather than renting someone else's. Which is most of what this talk is about.

ML and data at Colourbox
Thirty seconds, not three minutes. The only load-bearing part is the last card: everything in this talk follows from wanting to run it ourselves. The robotics background is worth one sentence because it explains the bias — coordination, autonomy and evaluation under uncertainty are the same problems wearing different clothes.
Setting the frame

Who here has tried agentic search?

Hands up.

Stop. Ask it, and actually wait for the hands — say the count out loud. Usually only a scattering go up, because "agentic search" sounds like something you would have had to set up on purpose. Do not explain anything yet. Nothing else goes on this slide; the whole point is that they answer before they know where it is going. Then advance.
Setting the frame

You are all already using one

When Claude Code answers a question about your repository, it is not reading the repository. It plans a search, runs it, opens what looked promising, and goes again.

// where does this thing retry?
grep -rn "retry" src/
  src/http/client.ts:82   if (shouldRetry(res)) {
  src/http/backoff.ts:14  export function backoff(n) {

cat src/http/backoff.ts
  // …exponential, capped at 30s

grep -rn "backoff(" src/
// → enough. answer the question.
That is agentic search. No one set it up — it came with the tool.

Why it works so well there

A repository is text, and grep is exact, instant and free. The agent can afford to look twenty times, because every look costs nothing and returns precisely what it asked for.

And that is the whole pattern

Plan, look, read, look again. No embeddings, no index, no ranking — just a very good exact-match tool over a corpus made of tokens.

cursor, copilot, deep research — same loop

Now point the same loop at our corpus

grep on a JPEG returns nothing. A drawing has no tokens to match. A scanned lokalplan is a picture of text, and the text is not in the file. The tool this agent got for free does not exist for an asset library — visual retrieval has to be learned, indexed and served before the agent has anything to call at all.

ML and data at Colourbox
Land the reveal before anything else: the gap between the two shows of hands is the whole point — they have all been using agentic search, they just did not have a name for it. This is the on-ramp, and for this room it is the most important slide in the first ten minutes. Start from the thing they already believe and only then make it unfamiliar. Walk the transcript out loud. It plans a search, reads the two files that looked relevant, notices a second thing worth checking, searches again, and stops when it has enough. Nobody would call that a retrieval system, but it is exactly one — the agent's whole view of the repository is whatever grep handed back. The reason it works there is worth naming explicitly, because it is the thing we do not get: exact match over text is free, instant and perfectly precise. There is no ranking problem, so there is no retrieval quality problem, so nobody has to think about any of this. Then the turn. Our corpus is images, drawings, scanned pages and video. There is no grep for a photograph. Everything the rest of this talk is about — embeddings, late interaction, the funnel, the vector store — exists to build the tool that this agent already had for free.
Why agents amplify retrieval failure

Your agents are only as good as your retrieval

Agentic search is a loop — plan, retrieve, reason, act — run many times over one question. Retrieve is the only step that gates what the model ever gets to see.

CONTEXT WINDOW every retrieved document lands here — and is re-read at every later step one bad result enters here… …and never leaves PLAN RETRIEVE REASON ACT and the loop runs again — a median of 24 search calls for a single question

A bad result doesn't visit. It moves in.

A false positive is not a wasted slot on a page of results that the user ignores. It enters the context, gets re-read at every later step, and shapes the next action. Retrieval error does not just persist — it compounds.

Which is why the rest of this is retrieval

Not because the agent part is uninteresting, but because it is the part we can least afford to get wrong.

dense embeddingsexact match late interactionthe funnel the vector storeour own eval set
ML and data at Colourbox
This slide gives the talk its direction, so do not rush it. Everything after it serves the claim in the headline, which is Atita Arora's phrase, from her opening talk at Context Camp — worth crediting out loud by name, it costs one sentence. Hugo Bowne-Anderson's write-up is where the slides and the recording are. Walk the diagram once, slowly, in this order: the loop runs, retrieve is the only box that touches the corpus, and everything it returns is written upward into the context window. Then the point: nothing is ever removed from up there. A false positive is not a bad result the user scrolls past — it is a document the model re-reads on every subsequent step and reasons from as if it were evidence. The number is worth saying precisely. A median of 24 search calls per question is GPT-5 paired with BM25 on BrowseComp-Plus, ninetieth percentile 35, maximum 63. The retriever matters to that count: a weaker first stage makes the agent search more, so bad retrieval costs latency and tokens before it costs correctness. The compounding is the part people underestimate, and it is worse for us than for a chat product, because our agents will have tools. A wrong document does not only produce a wrong sentence — it produces a wrong action, on a customer's own asset library. The chips are the map for the next half hour, and they regroup into the four preconditions on the closing slide. The two papers that argue the cost question from opposite directions come straight after this — the next two slides — so point forward, not to the end.
The default: one vector each

One vector for the query, one per asset

Before we spend the next half hour on it — this is what "dense search" means: two encoders, one pooled vector each, and a single similarity score at the end.

Encoded independently

Query and document never meet during encoding — so the whole document side is prepared offline, and query time is one similarity search.

But it makes one strong bet

That pool step preserved every clue a future query might care about — decided before anyone knew what would be asked.

ML and data at Colourbox
Keep this tight — it is a shared-vocabulary slide, not an argument, and for a room that has heard "embeddings" a hundred times without ever seeing the shape of one, thirty seconds here pays for the next twenty minutes. Walk the two lanes once: the query is split into tokens, each token gets a vector, and then pool crushes them into a single point. Same on the document side, offline. At query time the whole comparison is one dot product between two points. The last card is the hinge for the next slide: that pool step had to decide what was worth keeping before anyone had asked a question.
Agentic search · cost

Tied on quality, 1,431× apart on cost

Agents don't search once. When a workflow fans out fifty queries where a person issued one, cost per query stops being a rounding error and becomes the architecture.

Cost versus MTEB(LLM) score across 36 models: embedding models occupy the cheap end of the Pareto frontier, LLMs sit one to three orders of magnitude further right
Figure 1 from El Assadi, Muennighoff & Lee, The Embedder's Dilemma · COLM 2026. Note the log scale on cost.
0.4

points apart

Best LLM (Gemini 3.1 Pro, 77.6) against best embedding model (77.2), over 37 tasks. In aggregate, tied.

1,431×

the cost of closing it

USD 154 against USD 0.11 per benchmark pass — and 2.5–736× slower on the same GPU.

ML and data at Colourbox
Three things worth saying out loud beyond the numbers on screen. First, where it flips: LLMs genuinely lead on reasoning-heavy retrieval, embedding models lead on classification, and the two match on clustering, STS and pair classification — so this is a routing decision, not a winner. Second, reasoning tokens are 28–81% of LLM inference cost, and in their ablation lower reasoning budgets preserved or even improved retrieval quality — which is a knob we control because we host the serving stack. Third, the Pareto frontier holds the leading embedding models plus exactly one LLM, and that LLM buys 0.4 points for three orders of magnitude. For an agentic caller issuing many queries per task, that trade is not close.
Agentic search · the counter-example

And then someone trained a small one for exactly this job

SID-1 is an open 14B base, RL-trained end to end for one job — agentic retrieval. On its own benchmark it out-recalls the frontier models it costs a fraction of.

Recall against cost per question on 191 multi-hop questions: SID-1 sits above every frontier model at a fraction of the cost per question
191 multi-hop questions · 0.84 four-rollout, 0.78 single-pass · SID-1 technical report, 2025.
24×

faster, and far cheaper

5.5 seconds per question against GPT-5.1's 131 — and roughly 400× less per question than Sonnet 4.5. The Embedder's Dilemma, from the other side.

Trained for one job

Qwen3-14B, tuned with GRPO and no supervised warm-up on synthetic multi-hop questions.

Our read

The previous slide from the other end: the frontier model is better at everything in general and worse at this in particular, because the small one was trained on the retrieval task that actually gets run. Which is the argument for owning our search rather than renting it.

ML and data at Colourbox
Put this straight after the embedder's dilemma, because it is the same graph read backwards. There, closing a 0.4-point gap cost 1,431×. Here, a 14B model closes the gap in the other direction by being trained on the actual task. Be honest about the caveats or someone will find them. The 0.84 is the 4x setting — four rollouts in parallel at raised temperature, fused with reciprocal rank fusion — so it is buying some of that recall with test-time compute. The benchmark is their own 191 questions, half public and half written in-house, which is exactly the kind of home-ground advantage we would also have. And SID-1 itself is waitlisted, not downloadable. The number to be careful with is nDCG. The headline claim here is recall, where the gap is clear. On nDCG the two are much closer — close enough that "outperforms the frontier" would be overstating it — so quote recall, and read the table again before quoting nDCG at all. None of which weakens the point for us. The base model is open, the training method is published, and the thing that made it win was domain fit rather than scale. That is the most concrete evidence in the deck that a self-hosted retrieval model is a real option rather than a principled preference. If asked "so should we train one?" — not yet. The honest sequence is the eval set first, then measure how far a good off-the-shelf embedder gets us, and only then consider training. This slide is a reason to keep the option open, not a plan.
Part one

Search in our systems

Two products that look alike from the outside, with opposite priorities underneath.

The landscape

Two products, two retrieval problems

A stock library and a DAM system look similar from the outside. The search problem underneath is not the same one.

Colourbox

Stock photo library

One shared catalogue, many unrelated buyers. Every user searches the same corpus with no prior relationship to any of it. Intent is broad and visual — "two people laughing in an office, lots of copy space".

Millions of assets, one index, and no boundary between users — everybody searches the same catalogue.

Skyfish

Digital Asset Management (DAM) system

One customer's own assets, and they know what's in there. The user is often looking for a specific known item — a campaign shoot, a signed contract, last year's product video.

Every query is scoped to one tenant before it ranks — the boundary is part of the retrieval problem, not a detail around it.

ML and data at Colourbox
Set up the two shapes here — one shared catalogue against one tenant's own material — and leave recall versus precision to the next slide, which is where the architectural conclusion gets drawn. Do not make the same point twice.
Products one and two

Recall on one side, precision on the other — over six kinds of file

One search box each — and underneath, the priorities point in opposite directions.

product one · stock

Browse, don't locate

Recall over precision. Twenty good options beat one perfect one — the buyer is choosing on taste we cannot model.

The query is visual, not lexical. Mood, composition, colour, copy space — none of which anyone reliably typed into a metadata field.

And whole-asset relevance is enough. A photograph has no page 74. The compression that would ruin a precise lookup costs almost nothing when the target is a whole region of the space.

product two · DAM

Find the right one

Precision over recall. The user is after a specific known item, so a near miss is a failure rather than a compromise.

The corpus is mixed, and mostly visual. Images, video, PDF pages and a long tail of everything else — over one customer's own material.

And the answer is somewhere inside the file. A hundred-page report where the figure is on page 74; a two-hour inspection video where the moment is at 41 minutes. Deciding a document is relevant is not the same as pointing at the right place inside it — and part two is about scoring that can still point.

Kommuneplaner

geoimagetext
24,0 m

Lokalplaner

drawingpdftext

Historiske arkiver

textscanocr
12

Teknisk dokumentation

drawingpdftable

Produkt- og feltbilleder

imageexifcaption

Inspektions- og mødevideo

videoframestranscript

What one DAM customer calls "our files": six things, six different failure modes — and only one of them is solved by reading a caption.

ML and data at Colourbox
This is the comparison the rest of the talk keeps referring back to, so make it land: same search box, opposite priorities. Stock wants twenty good answers and can afford compression; the DAM wants the one right answer and cannot. The last line of the DAM card is the one that sets up part two, and it is deliberately mechanism-free — no vocabulary the room has not met yet. Say the page-74 example out loud and leave it hanging; the heatmap slides later are the answer to it, and it is much more satisfying if the question was asked here first. Point at the tiles rather than reading them out. If the room is Danish these are recognisable on sight; if not, the shapes still carry it. The takeaway is modality heterogeneity — this is what a customer means by "our files". Two are worth an extra beat. PDFs are the surprise: page images rather than extracted text, so layout, tables and figures stay searchable, and it is where late interaction has the clearest case. Video is the other: sampled to frames, one asset becomes many vectors and the temporal question — which moment — becomes unavoidable. And the long tail is real: office files, audio, archives, formats one customer cares about deeply and nobody else has. That tail is why "one pipeline per file type" loses.
Part two

How retrieval actually works

Why filename search is still the baseline, what beats it, why most of a DAM has to be searched as an image — and what all of it costs to run.

The puzzle

Why do DAM users still search by filename?

If we have strong vision models and strong embedding models, why does every customer still type a filename into the box?

search  "Q3_campaign_final_v2"
browse  /Clients/Bergen/2024-06/

They are not being primitive. Exact match came back into fashion for a reason, and it is worth understanding which reason.

ML and data at Colourbox
Ask the room to guess before advancing. Most people answer "infra" — which is half right, and the other half is the interesting half.
The puzzle

Filename search is hard to beat — and not only on infra

The boring infra reasons

  • Zero infrastructure — it already ships with the product
  • No model to deploy
  • No index to maintain
  • Privacy by default

But also: why it sticks

  • Exact matching is extremely strong for filenames.
  • Granularity is non-negotiable — when you want Q3_campaign_final_v2.psd, you want exactly that file.
  • Filenames and folder trees are structured data, and structure is cheap to exploit.

So the bar for a learned retriever is high

It has to earn its keep by adding semantics without losing exact local clues — and justify deploying a model and maintaining an index on top.

ML and data at Colourbox
The right-hand column is the part people underrate. Spend your time there.
Is exact match all we need?

Exact match is a gift — only when the vocabulary lines up

The wrong conclusion would be that BM25 is all we need.

BM25 performance collapsing on the synonym variant of the LIMIT corpus
Figure: Amélie Chatelain, Late Interaction Field Guide, from data in Weller et al., LIMIT

Vocabulary mismatch is alive and well

Exact retrieval breaks the moment the user says "the beach shoot" and the asset is called DSC02074.JPG, in /Uploads/2024-06/, description empty.

On the synonym variant of LIMIT — a benchmark where documents carry combinations of attributes and queries ask for those combinations — BM25 collapses. And in a DAM the uncurated folder is the default state.

So we want both

Exact-ish local evidence, which lexical search gives us free — and soft semantic matching, which dense retrieval gives us free.

At the same time. That is the whole ask.

ML and data at Colourbox
Be explicit that this is not a pro-lexical slide. It sets up the "we want both" framing that late interaction answers.
Something on this chart

Surely we can just increase the dimensions?

Retrieval quality against embedding dimension — and one line here does not behave like the other dense models.

Retrieval quality against embedding dimension, with one line behaving unlike the other dense models
Figure: Amélie Chatelain, Multi-vector Search: A Late Interaction Field Guide
GTE-ModernColBERT

One line here does not behave like the others.

Is it just a huge pooled vector?

No. It uses a different retrieval paradigm.

So a wider vector is not the fix. Which leaves the real question: what exactly gets lost when we pool one?

ML and data at Colourbox
This is the hinge of the talk. Pause before turning the page.
The bottleneck

One vector has to carry every clue

Not because the vector is too narrow — because a single query can hold several independent constraints, and pooling has to choose between them.

"two people laughing in a bright office, wide shot, copy space on the left"

2 peoplelaughing officewide shot copy space left

↓ compressed into

[ single vector · dim d ]

→ diluted signal

"Copy space on the left" is a fact about a region of the image. Pooling decided what to keep before it knew anyone would ask.

So what if we gave every token its own vector?

Five constraints, five query vectors — and the asset keeping one vector per unit of itself rather than one for the whole thing. Nothing gets averaged, so "copy space on the left" stays its own piece of evidence, free to be matched by the part of the image that actually has copy space on the left.

ML and data at Colourbox
Placed here on purpose. The chart before it showed that width is not the fix; this is why — the constraints in one query are independent of each other, so pooling loses some of them no matter how many dimensions it has to play with. This is not "a semantic question about authentication" — it is five constraints that all have to survive. Read the query out and count them on your fingers. Then pose the last card as a genuine question rather than reading it out — the room can usually see the move once the five chips are on screen. Do not name it. "One vector per token" is the whole idea, and the next slide is where it gets its name, its three-way comparison and its cost. This is the bottleneck to keep referring back to.
Escaping the ceiling

Late interaction sits in the middle

Three ways to score a query against a document. The difference is when the two sides are allowed to meet.

← cheaper · more scalable more accurate · expensive →

Single-vector retrievalpooled dual encoder

QUERYDOCUMENTqdq1q2q3q4d1d2d3d4d5d6EncoderEncoderpoolpoolone dot productpool(q) · pool(d)
docs encoded offline

Late-interaction retrievalmulti-vector dual encoder

QUERYDOCUMENTqdq1q2q3q4d1d2d3d4d5d6EncoderEncodersum of best matchesΣ max · nothing pooled
docs encoded offline

Full-interaction rerankingcross-encoder

QUERYDOCUMENTqdq1q2q3q4d1d2d3d4d5d6Encoderone score
query-time only

That interaction has a name: MaxSim

For every query token, keep only its best-matching unit of the asset, then sum those maxima. Nothing is pooled, so a rare local clue survives.

And for us the unit is a patch

Same operator, patches instead of tokens — which is where the next few slides go.

ML and data at Colourbox
Read the three panels left to right, and read them as the same picture drawn three times — the only thing that changes is where the query and the document are allowed to meet. Left: both sides are encoded separately, then each side is squeezed through pool down to one vector, and the score is a single dot product. Cheap, scalable, and the compression happens before anyone knows what will be asked. Middle: same two encoders, same offline document side — but nothing is pooled. Every query vector is compared against every document vector, and only the best match per query vector counts. The four bold lines are those maxima; the faint ones are the comparisons that lose. Right: there is only one encoder, and the query and the document go into it together. Every token attends to every token, which is why it scores best and why nothing can be precomputed — the badge is the whole argument, "query-time only" means you can never run this over the catalogue, only over a shortlist. The key words are "don't pool" and "delay". Everything deployable about the middle column comes from documents still being encoded offline. Define MaxSim carefully here — this is where the thing we have been calling a similarity score gets its name, and every later slide leans on it. The patch card is a promise, not a new idea: say "the unit is whatever we chose to keep vectors for", and that the next few slides show it on a real page. Do not explain the geometry here — the ColPali slide does that. Walk one query token: it looks at every unit of the asset, keeps only its best match, and contributes that one number. Sum over query tokens and that is the score. The example worth saying out loud is the multilingual one: the customer types English, the asset was captioned in Norwegian. Exact match scores zero there, and a pooled vector has already blurred it. MaxSim can still find the one unit that carries the meaning. The patch line is the bridge to the DAM. A page is an image with structure, so nothing about the operator has to change when we move from captions to documents.
Why visual retrieval at all · stock

A picture is worth a thousand keywords

Sometimes the evidence you are looking for is a clean line of text. Sometimes it is simply not.

Hieronymus Bosch, The Garden of Earthly Delights: a dense triptych with hundreds of small independent scenes
Figure: Amélie Chatelain, Late Interaction Field Guide · The Garden of Earthly Delights, H. Bosch, Museo del Prado · example from Benjamin Clavié's talks
the query

"painting with a guy stuck in a mussel"

What a caption pipeline gives us

No text on the asset at all. To match this query the pipeline would have to have guessed, at indexing time, that this one odd detail would matter — or be exhaustively descriptive about everything.

This is our contributor metadata problem

A stock contributor types eight keywords. A DAM customer types none. Neither of them anticipated the query — and a caption is a lossy, one-shot bet made before the query exists.

ML and data at Colourbox
Read the query out loud, then let people hunt for it in the painting. It takes a while, and that is the point. A traditional text pipeline cannot touch an image without a captioning model in front of it. And captioning is exactly where this breaks: the model would have to decide in advance that a man stuck inside a mussel shell is the salient fact about this painting. It will not. It will say "surreal triptych with many figures". Bring it home: this is our metadata situation with the labels changed. Contributor keywords on stock, near-nothing on a new DAM tenant. Direct visual retrieval keeps regions of the asset available for matching instead of forcing everything through a text bottleneck first.
Why visual retrieval at all

And in a DAM, the picture is usually a page

Paintings are a fun example. A customer's own report PDF is the case we are actually paid to solve.

A report page: choropleth map of EU member states shaded by whether they have implemented Individual Learning Accounts, with a legend and a short caption
Figure: Amélie Chatelain, Late Interaction Field Guide · page from Loison et al., ViDoRe v3
the query

"is there uniform interest across all EU regions in adopting Individual Learning Accounts?"

What OCR gives us

The caption, if we are lucky, and a handful of country labels. The answer lives in the spatial pattern of the shading — which no OCR pass and no generic caption preserves.

Preprocessing is not neutral

Whatever OCR or captioning fails to keep is gone before ranking starts. That is an irreversible decision taken at index time, by a component nobody thinks of as part of the ranker.

ML and data at Colourbox
The natural objection to the painting slide is "I'm not retrieving paintings". This is the answer. A very large share of enterprise documents carry their meaning in figures, and our DAM tenants upload exactly this kind of material. The last card is the one worth pausing on. We tend to treat OCR and captioning as plumbing, upstream of the interesting part. But they decide what the ranker is even allowed to see. A pipeline that reads "map of Europe" off this page has already lost, no matter how good the retriever behind it is. Note in passing that this happens to be an EU policy report — the kind of document a public-sector tenant would want indexed without it leaving European infrastructure.
Going multi-vector · ColPali

From tokens to patches, on a page

The operator we already built. Cut the page into a grid of patches, keep one vector per patch — and we can then see where the score came from.

The CDER bar chart with a regular patch grid drawn over it, so the title, the legend, the axis labels and individual bars fall into different cells
Figure: Amélie Chatelain, Late Interaction Field Guide · after Faysse et al., ColPali

Same idea, new geometry

Text keeps token vectors along a sequence; a page keeps them across a two-dimensional layout. Pages, photos and frames are all patch grids — one pipeline.

so let's ask one

"is there a rise in CDER NME submissions from 2007 to 2008?"

CDERNME 20072008

Which regions should it notice? Three unrelated parts of one page.

And we can show our work

Scoring is per query token against per patch, so the match can be highlighted on the page.

ML and data at Colourbox
MaxSim is already on the table from the late-interaction slide, so use the name freely here and spend the time on what is new. This is that same operator on a real page — the only genuinely new thing is the geometry: a sequence of token vectors becomes a grid of patch vectors. Same encoder-offline, same scoring, same reason it works — local evidence stays independently addressable. The first card carries the strategic payoff and is worth saying explicitly, because it is what justifies the index cost: one representation for images, PDF pages and video frames means one pipeline instead of four. Then make the room commit before advancing. Read the query out, ask which regions of the page they would want the retriever to look at, and let someone actually answer. The next two slides show the real heatmaps one token at a time, and the payoff is much better if people have guessed first. The explainability point in the dark card is not decoration — DAM customers ask for it, and a single pooled vector fundamentally cannot give it to them.
Interpretability

The token 2008 lands on the axis

Best-matching patches for one query token. The strongest single cell is the 2008 tick label, with a warm band along the whole x-axis.

The same chart desaturated, with a similarity heatmap overlaid: the brightest region is the 2008 tick label on the x-axis, with a softer glow along the axis and over the mid bars
Figure: Amélie Chatelain, Late Interaction Field Guide · after Faysse et al., ColPali

One row of the MaxSim matrix

This is literally what the operator does: for the query token 2008, score every patch, keep the best one, contribute it to the sum.

It found the right kind of place

Not just the exact label — the surrounding axis region lights up too. Soft semantic matching and exact local evidence, in the same score.

Nobody trained it to do this

Localisation falls out of never pooling the patch vectors — a side effect of the architecture.

ML and data at Colourbox
One beat per token. Here: the model has found the 2008 tick label, and it has also warmed the rest of the axis, which is the right neighbourhood. Say the quiet part: nobody trained this model on "highlight the axis". The localisation falls out of keeping one vector per patch and never pooling them. Interpretability here is a free side effect of the architecture, not a feature someone bolted on.
From evidence to production

So why isn't this everywhere yet?

The evidence is strong and the operator is simple. What stops it is engineering — three objections: one we simply pay, two with real answers.

01

Storage

One vector per patch, across the whole catalogue.

Conceded. This is the bill for keeping local evidence. Quantization softens it; nothing removes it.

02

Speed

Scoring is no longer one dot product per candidate.

03

Ecosystem

Which model, on which engine, and can we still swap it later?

Breath slide. Name the three objections, and concede storage out loud rather than leaving it hanging — it is on the card now, so say it and move on. The next slides answer speed and ecosystem, which are the two with real answers.
The speed objection

Naive MaxSim is doomed

Compare every query token against every unit of every asset and the arithmetic ends the discussion before we get to latency.

All-to-all comparison: every query token connected to every document token of one document
Figure: Amélie Chatelain, Late Interaction Field Guide. Counts are her MS MARCO worked example.
~32

query tokens

Every query token needs its own best match on the asset side.

~80

tokens per passage

All token pairs get scored inside a single document.

~9M

passages

Then repeat that all-to-all comparison across the corpus.

~23B

dot products, per query

At dimension 128. Not a first-stage plan under any latency budget.

ML and data at Colourbox
We already conceded that late interaction costs index size and scoring time. This is where we answer how it is made to work. Start with the honest worst case. If you take MaxSim literally — every query token against every token of every document — you get roughly 32 × 80 × 9M dot products at dimension 128 for MS MARCO. About 23 billion. That is not slow, that is impossible. The important reframe: no serious system implements exact MaxSim over the whole corpus. MaxSim is the score we are trying to arrive at, not the plan for getting there. Everything in the next few slides is about reaching a good candidate set without paying the target cost everywhere. If someone asks what it looks like for us, the shape is the same with patches instead of tokens and one tenant's assets instead of passages — same verdict, an order of magnitude down. Say it, do not put it on the slide: the MS MARCO example already makes the point, and our own count is the one number in it nobody can check from the room.
Making it work

Every fast system is a funnel

Two things change on the way down: the set shrinks, the price per asset rises.

THE SET · SHRINKS ↓ Every asset in the index cheap score · nothing decompressed Candidates gathered from the cheap structure Shortlist approximate MaxSim, then prune Survivors exact MaxSim

Stage widths left blank on purpose — they come out of measuring recall against candidate count on our own corpus. The shape of the funnel is the point; the numbers are a measurement we have not taken yet.

The scorer · cost per asset rises ↓

1 · Candidate generation cheap

Score every query token against centroids in the approximate index.

2 · Approximate MaxSim, then prune moderate

Score documents as bags of centroids, nothing decompressed yet.

3 · Exact MaxSim expensive

Run only on the survivors — where we are willing to pay full price per asset.

And you don't build this funnel — you configure it. Every stage is already a feature of a modern vector database: an approximate index like HNSW for stage 1, sparse keys and tenant filters for stage 2, a MaxSim comparator for stage 3.

ML and data at Colourbox
This is the pattern, and it is not exotic — it is the same funnel we already use everywhere else in search. Start broad and cheap, spend expensive scoring only on a small set. Read it top to bottom: every asset gets a coarse score, we gather candidates, we prune with an approximate MaxSim, and only the survivors get the real thing. As the set shrinks we can afford more per asset. So the knob is not mysterious. It is: how many assets survive long enough to receive exact scoring? That single number is what we tune against our latency budget. I have deliberately left the stage widths blank. Her illustrative numbers were 9M → 100k → 1k → top-10 on MS MARCO; ours have to come from measuring recall against candidate count on our own corpus, and I would rather show nothing than show a number I made up. The last card matters for the next slide, and it is the relief beat: this sounds like a research architecture and it is actually a config file. Map the three stages onto the three database features out loud — same order, same words. If the index-versus-database distinction comes up, that is the answer to give: an index does the maths, a database is what makes it operable — filters, persistence, and add, update and delete without a rebuild, which a DAM needs all day. The filtering line deserves one beat if anyone bites. Post-filtering is the naive implementation and it silently destroys recall: you ask for the hundred nearest and then throw away the ninety-eight that belong to other tenants. Everything interesting in these systems is about keeping a filtered approximate search fast. If someone asks "why not just pgvector" — completely reasonable at small scale, and one fewer system to run. It gets harder when you want multivectors and quantization on a corpus this size.
Infrastructure · the vector store

Where the vectors actually live

For us that database is Qdrant, self-hosted — one engine holding all three retrieval arms, which is what keeps a mutable index consistent with the assets.

AssetsTritonembedding + vision modelsembedupsert vectorsQdrantself-hosted, on our own nodesPREFETCHDense · HNSWsemantic similarity, payload-filteredSparse · BM25exact tokens: filenames, product codesLate interactionmultivectors, MaxSim over patchesone query, three arms in one requestfuseRerankMaxSim on the survivorsthe expensive score, lasta user questionAgentLangChain · deepagentsprompt + hitsInference APIself-hosted, on vLLMAnswer · grounded in the hitstop-k hitssearchINDEX TIME · OFFLINE, ONCE PER ASSETQUERY TIME · ON EVERY QUESTION
ML and data at Colourbox
This is the slide that turns the previous three from an architecture lecture into something we operate. The funnel is not ours to invent — it is prefetch plus rerank in a query body. Walk the diagram in two passes. First the top: index time happens offline and once per asset — the embedding and vision models are served on Triton, and what comes out is upserted into Qdrant. Nothing there is on the critical path of a search. Then the inside of the box, which is the point of the slide. One query fans out to three arms in a single request: dense HNSW for semantic similarity with payload filters, sparse BM25 for the exact tokens a DAM user actually types — filenames, product codes — and late interaction over multivectors. They are fused, and only the survivors reach the rerank, where MaxSim is finally paid for. That is the funnel from two slides ago, expressed as one round trip instead of three services, and all three arms live in the same engine rather than in three stores we would have to keep consistent. Then the right-hand column, which is now split into the two things people conflate. The agent is application code — LangChain, deepagents — and it is the thing that decides to search, reads the hits and decides whether to search again. The model behind it is a separate service: our own inference API, served by vLLM on our own GPUs. The agent prompts it; it does not run inside it. That is the retrieval step from the very first diagram in this talk, with real component names on it. Worth saying explicitly, and then leaving alone until the closing slide: nothing in that picture is a service we rent. Triton, Qdrant and vLLM are three processes on our own nodes — Apache-2.0, Rust, one binary, and the company behind Qdrant is Berlin-based. Do not make the sovereignty argument here; it has its own slide at the end. The reason one engine matters more than raw benchmark speed: a DAM index is mutable all day, and every extra store is another thing to keep consistent with the assets, another backup story, another delete path to get right under GDPR. Two systems is a real operational cost and should be earned. Do not oversell Qdrant against the specialised engines. On pure multi-vector first-stage throughput a purpose-built engine wins, and if that ever becomes our bottleneck we should use one — for that tenant, not for the whole product. The Berlin point is worth one sentence and no more. It is a genuine reinforcement of part two, but the substantive claim was always about self-hosting rather than the vendor's postcode.
Part three

Data and evaluation

The half that decides whether any of the above works — and the half nobody is going to publish for us.

One piece of vocabulary first: an "eval set" is a test suite for search

A list of real queries, and for each one the asset that should come back — plus a number saying how often it did. Same idea as the tests you already write: the assertion is just "the right thing ranked first" instead of "the function returned 4". Everything in this part depends on having one.

Two sentences, then move. The only job here is that nobody in the room is still quietly wondering what an eval set is — the term has been on the chip list since the retrieval slide and it gets used constantly from here on. If the room is engineers, the test-suite analogy is the whole explanation: same discipline, same reason you cannot refactor without one, and the assertion is just fuzzier.
The starting point

The open data is good, and it isn't about us

Danish and its neighbours are served by a handful of genuinely good open projects. None of them can tell us whether search works on our customers' archives.

What does exist — and is good

  • Danish Foundation Models — an open Danish consortium building models and, just as importantly, curating the corpora behind them.
  • Danish Gigaword — a deliberately broad open Danish text corpus, spread across domains rather than scraped from one.
  • MTEB — the Scandinavian Embedding Benchmark has moved into it, so the Nordic retrieval tasks now live there. ScandEval shares much of the same pool.
  • The Nordic Pile and the National Library of Norway's NB AI-Lab — the same work, one border over.
credit where it is due

Where all of it stops short

  • It is prose. Text corpora and text benchmarks. Our assets are images, drawings, scanned PDFs and video.
  • It is general domain. Nothing in it resembles one company's naming conventions, product codes or internal shorthand.
  • It is not known-item retrieval. Almost no public benchmark measures "did the right asset come back first", which is the DAM question.

So the gap is not "Danish models". The gap is Danish retrieval evidence on documents shaped like our customers' — and nobody is going to publish that for us.

ML and data at Colourbox
Be generous about the Danish projects — several people in this room will know them, and the argument is stronger if it is clearly not a complaint about Danish NLP. The point is a modality and task gap, not a language-resource gap.
The way out

So we build the data ourselves

Representative beats large. The only corpus genuinely representative of our search problem is the one already sitting in our own storage.

  • Representative, not public. An eval drawn from our own assets measures the product. One drawn from a public benchmark measures the benchmark.
  • We are the ones allowed to. Our own systems and storage — plus data we collect from sources we have actually talked to. Not an indiscriminate crawl: a permissioned one, with an agreement behind each source, and nothing crossing a border to get labelled.
  • And it compounds. Every labelled query becomes an asset a competitor cannot download. The moat is the eval set, not the model.

But there is a prerequisite

You cannot sample representatively from a corpus you do not understand. Before any of this: what do we actually store, in which modality, for whom, and how much of it?

That is the unglamorous half of data sovereignty: owning it is the legal claim, knowing it like the back of your hand is what makes owning it worth anything.

the actual first step
ML and data at Colourbox
This is the hinge of part three. Everything after it — the modality mix, the model choice, the eval design — is downstream of knowing the corpus. Say the last line slowly.
The tradeoff without data

The trap in our own query logs

Our customers have only ever had keyword search. So that is what they type — and if we learn from it uncritically, we will carefully rebuild it.

1

They type what has worked

Filenames, product codes, exact tokens. Nobody types "two people laughing in an office" into a system that has never once rewarded it.

2

The logs inherit the habit

Sample an eval set from those queries and nearly every relevant judgement turns out to be satisfiable by exact match. The bias is now baked into the instrument.

3

The eval certifies the baseline

A semantic model scores worse on our own benchmark than the filename search it was meant to replace. The model is not wrong. The ruler is.

Which makes the model choice and the eval choice one decision

A keyword-shaped eval systematically flatters sparse and exact matching and systematically punishes dense semantic embeddings — so reading "dense versus sparse" off that eval is how you conclude, wrongly, that semantics does not help here. We come back to this when we pick models, because it is the same trap wearing a different hat.

ML and data at Colourbox
This is the slide that stops someone in the room saying "just A/B it on the logs". You cannot A/B your way out of a biased instrument. Mitigations worth mentioning if asked: seeded semantic queries written against the corpus, known-item tasks constructed from asset content rather than from logs, and interleaving rather than absolute comparison.
Back to infrastructure

Which is why we own the serving stack

A biased eval is fixable — but only if you can re-embed the corpus and measure again. That is what owning the stack buys.

01

Swap the model

The zoo expands monthly. A checkpoint we can pull, serve behind the same Triton endpoint and benchmark ourselves is a checkpoint we can replace — with no vendor roadmap in the way.

the zoo keeps moving
02

Re-index against our own eval set

An eval only means something if we can rebuild the index for a new candidate. Batch embedding on our own GPUs makes re-indexing a decision we take, not an invoice we receive.

measure, then commit
03

Never leave the jurisdiction

Every eval run, every re-index, every captioning pass is customer material going through a model. On our hardware in EU data centres it stays inside European jurisdiction throughout.

sovereignty, in practice

The trade-off stays ours

Quality against storage against latency against freshness — and it resolves differently per product. Stock takes the dense, cheap-index end; DAM pays for patches, because there a near miss is a failure. Owning the stack is what lets us sit at two points on that curve at once — and change our mind when the next model lands.

All of it on our own GPUs — inside EU data centres.

ML and data at Colourbox
Close the loop back to the infrastructure argument from part two. The whole self-hosting argument was framed as sovereignty, and it is — but this is the second half of the return on it. Because we run the models, we can pull a new checkpoint on Monday, re-embed a sample of a customer's corpus on Tuesday and have a number on Wednesday. A hosted API gives you none of that: you cannot benchmark what you cannot re-index, and you cannot re-index a corpus you are not allowed to send. Sovereignty and the ability to keep improving turn out to be the same piece of infrastructure.
Where this goes

An agent searching a DAM
needs all of it

Not four separate projects. Four preconditions for the same thing — and a DAM is where they all bind at once: a customer's own assets, in every format, where a near miss is a failure.

It has to see

Content understanding over images, video, PDF pages and the long tail — because the evidence is in the asset, not the caption someone forgot to write.

patches, not metadata

It has to be right

Late interaction where precision decides the outcome, and an eval set built from our own customers' queries rather than a public leaderboard.

near miss = failure

It has to afford to ask

An agent fans out. The funnel, the compression and the embedding-first economics are what make the fiftieth query as cheap as the first.

cost is the design

It has to stay ours

Self-hosted on EU infrastructure, so we can swap the model, re-index on our own terms, and never move a customer's assets out of jurisdiction.

sovereign by architecture

Which leaves the interesting questions open

What does an agent need from search that a person never asked for — an explanation, a confidence, a second opinion? What does our eval set become when the caller is a machine that will happily ask a hundred times? And which of these do we build next?

Do not answer the three questions. This is the slide that should start the discussion, so stop talking and let the room take it. If it stalls, the one to pull on is the eval question — everyone in the room has an opinion about what "good" means for their own corpus, and that is exactly the conversation worth having.