ModelTencent HunyuanTencent Hunyuanpublished Sep 4, 2026seen 1d

tencent/EVIE-4.5B

Open original ↗

Captured source

source ↗
published Sep 4, 2026seen 1dcaptured 1dhttp 200method plaintask visual-document-retrievallicense apache-2.0library colpali-engineparams 4.5Bdownloads 497likes 16

---

> 📢 Release Announcement: All model weights, training pipelines, token compression algorithms (HAC), and evaluation suites have been fully open-sourced in this repository. Full technical details, architectural ablations, and the formal research paper will be updated in an upcoming release.

---

🌟 Highlights

  • Top-Tier Benchmark Performance: 66.75 on ViDoRe V3 for EVIE-8B and 66.02 for EVIE-4.5B with single-projection Prefix-MRL.
  • ⚡ Prefix-MRL Elasticity: Single 2048D linear projection. Freely truncate at runtime into $\{64, 128, 256, 512, 1024, 2048\}$ dimensions without separate models.
  • 📦 Ultra-Compact Index (HAC): Training-free Hierarchical Agglomerative Clustering compresses token counts from ~750 down to 32 vectors/page, slashing index storage to 3.81 GiB per million pages.
  • 🌐 138 Multilingual Tasks Evaluated: Thoroughly evaluated across ViDoRe V1, V2, V3, and JinaVDR across 4 metric families (nDCG, Recall, MAP, MRR @1/5/10).
  • 🔬 ARD Distillation Recipe: Anchor-preserving, capacity-aware relation distillation reproducing full student training from the 8B teacher.

---

🧠 Architecture & Technical Highlights

Query Text ────────► ColQwen3.5 (BiDir Attention) ────► Elastic Multi-Vectors (64D–2048D)
│
MaxSim Matching
│
Doc Image ────────► ColQwen3.5 (Vision Encoder) ────► HAC Compression ──► 32 Vectors / Page
  • Late-Interaction Multi-Vector Paradigm: Unlike dense single-vector retrieval that collapses high-resolution document pages into a single point, EVIE preserves fine-grained visual details (complex tables, layout structures, charts, and small typography) through token-level representations, scoring relevance via late-interaction MaxSim:

$$ S(Q, D) = \sum_{i=1}^{|Q|} \max_{j=1}^{|D|} (q_i \cdot d_j) $$

  • Prefix-MRL (Single-Head Elastic Representation): EVIE-4.5B introduces single-projection Prefix-MRL. A single 2048D linear projection natively supports runtime truncation down to {64, 128, 256, 512, 1024, 2048} dimensions without maintaining multiple heads or separate checkpoints.
  • ARD (Anchor-preserving Relation Distillation): The 4.5B student is distilled from the 8B teacher using token-relation topological geometry, hard-negative margin calibration, and anchor-preserving alignment, maintaining peak retrieval accuracy even under low-dimensional prefixes.
  • HAC Token Compression (Hierarchical Agglomerative Clustering): A plug-and-play, training-free token reduction algorithm that aggregates visual patch tokens into 32 or 64 semantic centroids in joint feature-spatial space, reducing 1M-page index footprints to as little as 3.81 GiB.

---

📊 Comprehensive ViDoRe Leaderboard Comparison

Performance comparison across modern multi-vector late-interaction visual document retrievers on ViDoRe:

| Rank | Model | Base Model | Param | Embed Dim | ViDoRe V1 (nDCG@5) | ViDoRe V2 (nDCG@5) | ViDoRe V3 (nDCG@10) | | :---: | :--- | :---: | :---: | :---: | :---: | :---: | :---: | | 🥇 | [EVIE-8B](https://huggingface.co/tencent/EVIE-8B) | Qwen3.5-9B | 8.41B | 4096D | 92.18 | 74.23 | 66.75 | | 🥈 | [EVIE-4.5B](https://huggingface.co/tencent/EVIE-4.5B) | Qwen3.5-4B | 4.61B | 64–2048D Prefix-MRL | 92.07 | 73.38 | 66.02 | | 🥉 | EVIE-Preview-4.5B | Qwen3.5-4B | 4.54B | 128D | 91.73 | 70.87 | 65.36 | | 4 | webAI-ColVec1.1-8b | Qwen2.5-VL | 8.40B | 640D | 91.30 | 65.82 | 65.32 | | 5 | VultronRetrieverPrime-8B | Qwen3.5-9B | 8.40B | 320D | 92.08 | 68.18 | 64.26 | | 6 | webAI-ColVec1.1-4b | Qwen2.5-VL | 4.54B | 640D | 90.49 | 63.60 | 63.90 | | 7 | VultronRetrieverCore-4.5B | Qwen3.5-4B | 4.50B | 320D | 92.21 | 66.12 | 63.57 | | 8 | nemotron-colembed-vl-8b-v2 | Nemotron-8B | 8.80B | 4096D | 92.65 | 65.16 | 63.54 | | 9 | tomoro-colqwen3-embed-8b | Qwen2.5-VL | 8.00B | 320D | 90.76 | 65.40 | 61.60 | | 10 | nemotron-colembed-vl-4b-v2 | Nemotron-4B | 4.80B | 2560D | 91.62 | 64.49 | 61.42 | | 11 | athrael-soju/colqwen3.5-4.5B-v3 | Qwen3.5-4B | 4.60B | 128D | 91.54 | 64.25 | 61.46 | | 12 | tomoro-colqwen3-embed-4b | Qwen2.5-VL | 4.00B | 320D | 90.57 | 64.69 | 60.16 | | 13 | VultronRetrieverFlash-0.8B | Qwen3.5-0.8B | 0.85B | 320D | 88.15 | 60.36 | 56.16 |

---

🔍 ViDoRe V3 Per-Domain Breakdown (nDCG@10)

| Model | Avg | CompSci | Energy | Finance EN | Finance FR | HR | Industrial | Pharma | Physics | | :--- | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: | | EVIE-8B | 66.75 | 81.86 | 72.51 | 71.23 | 56.40 | 69.29 | 59.77 | 70.81 | 52.11 | | EVIE-4.5B | 66.02 | 81.72 | 72.32 | 70.00 | 54.90 | 67.82 | 59.40 | 70.27 | 51.69 | | webAI-ColVec1.1-8b | 65.32 | 80.08 | 70.12 | 71.90 | 54.87 | 68.55 | 57.65 | 67.88 | 51.50 | | nemotron-colembed-vl-8b-v2 | 63.54 | 79.30 | 69.82 | 67.29 | 51.54 | 66.32 | 56.03 | 67.19 | 50.84 | | VultronRetrieverPrime-8B | 64.26 | 79.80 | 70.30 | 69.00 | 54.50 | 66.80 | 57.40 | 68.20 | 51.70 | | VultronRetrieverCore-4.5B | 63.57 | 79.80 | 69.20 | 68.90 | 52.00 | 66.10 | 56.10 | 67.50 | 50.20 | | tomoro-colqwen3-embed-8b | 61.60 | 75.35 | 68.41 | 65.08 | 49.10 | 63.98 | 54.41 | 66.36 | 50.13 |

---

🎯 Prefix-MRL Elastic Multi-Vector Head

EVIE-4.5B embeds document and query tokens with a single 2048D linear projection head trained via ARD. You can truncate the channel dimension on-the-fly without maintaining different models:

Full Projection (2048D) [========================================================] 66.02
Prefix 1024D [============================] 65.94
Prefix 512D [==============] 65.90
Prefix 256D [=======] 65.68
Prefix 128D [===] 65.27
Prefix 64D [=] 64.51

| Dimension | Bytes / Vector | ViDoRe V1 | ViDoRe V2 | ViDoRe V3 | JinaVDR | 138-Task Avg4 | | :---: | :---: | :---: | :---: | :---: | :---: | :---: | | 64 | 128 B | 92.16 | 73.18 | 64.51 | 81.00 | 77.71 | | 128 | 256 B | 92.28 | 73.37 | 65.27 | 81.83 | 78.18 | | 256 | 512 B | 92.20 | 73.83 | 65.68 | 82.09 | 78.45 | | 512 | 1 KiB | 92.38 | 74.53 | 65.90 | 82.27 | 78.77 | | 1024 | 2 KiB | 92.39 | 74.53 | 65.94 | 82.43 | 78.82 | | 2048 | 4 KiB | 92.53 | 74.91 | 66.02 | 82.48 | 78.98 |

---

🗜️ Token Compression (HAC)

Raw late-interaction representations keep all visual patch vectors (~750 vectors/page), requiring substantial storage. EVIE integrates Hierarchical Agglomerative Clustering (HAC) in a joint semantic-position space...

Excerpt shown — open the source for the full document.

Notability

notability 7.0/10

Tencent Hunyuan released 4.5B model, notable but not frontier

Tencent Hunyuan has a model signal matching data demand, evals and quality.