AI & ML Infrastructure

Vision-language models and multimodal RAG on compressed visual tokens

Polygen compresses the visual tokens a vision-language model attends over. On Qwen2-VL-7B the recommended setting delivers +8.4 points of ScienceQA accuracy over the un-spliced baseline with 41% fewer attention FLOPs, and a visual corpus encoded once is 5.5x smaller to store and answers the same questions.
Business Impact

What changes when raw data stops moving

Raise accuracy and cut attention compute at the same time
Store multimodal corpora as descriptors, not cached embeddings
Transfer one small adapter across benchmarks without retraining
Swap one wrapper for the model class; keep the rest of the stack

+8.4pp

ScienceQA accuracy over baseline
Qwen2-VL-7B at T=3, n=200, NVIDIA H200, measured

~71%

Fewer attention FLOPs
At T=2 with +7.4pp accuracy; a FLOP count, not wall-clock

5.5x

Smaller RAG index
10.4 MB against 56.7 MB cached fp16 embeddings, 8-page corpus, measured

~3.7x

Requests per GPU on video frames
About half the visual tokens at the T=2 bypass point, measured
The challenge

Visual tokens are the cost centre of a vision-language model

A VLM feeds hundreds of visual tokens per photo, and thousands per document page, into an LLM whose attention cost grows with the square of sequence length. Multimodal RAG then caches fp16 embeddings for every page it might retrieve. Both bills scale with the token count, and neither improves the answer.
Approach

How Datasent enables this use case

Splice

Compress at the merger output

Polygen replaces the per-image visual token sequence with a much shorter structured representation, spliced where the vision encoder hands off to the LLM. Five operating tiers trade compression against fidelity; T=4 is lossless storage.
Adapt

One small LoRA, trained once

A rank-16 adapter trained on 300 ScienceQA images teaches the LLM to read the compressed sequence. It transfers to MMMU, VQAv2 and DocVQA without retraining; every lift on those benchmarks is cross-benchmark transfer.
Serve

Encode once, query many

A corpus is encoded once and stored as polygen descriptors. Each query runs against the compressed representation; the receiver reconstructs what it needs from tokens alone. Single-shot wall-clock savings are not claimed; the measured wins are accuracy, attention FLOPs, storage and requests per GPU.