Splice
Compress at the merger output
Polygen replaces the per-image visual token sequence with a much shorter structured representation, spliced where the vision encoder hands off to the LLM. Five operating tiers trade compression against fidelity; T=4 is lossless storage.

Adapt
One small LoRA, trained once
A rank-16 adapter trained on 300 ScienceQA images teaches the LLM to read the compressed sequence. It transfers to MMMU, VQAv2 and DocVQA without retraining; every lift on those benchmarks is cross-benchmark transfer.

Serve
Encode once, query many
A corpus is encoded once and stored as polygen descriptors. Each query runs against the compressed representation; the receiver reconstructs what it needs from tokens alone. Single-shot wall-clock savings are not claimed; the measured wins are accuracy, attention FLOPs, storage and requests per GPU.