We lifted 3 functions out of this paper's own repositories and ran 2 of them in a sandbox. "Ran" means the function executed on a synthesized input and returned a value. It is not a reproduction of the paper's results.
| Repository | Role | Ran |
|---|---|---|
| illuin-tech/vidore-benchmark | canonical | 2 of 3 |
| Function | Status | Where it lives |
|---|---|---|
| aggregate_results | Ran | illuin-tech/vidore-benchmark/src/vidore_benchmark/pipeline_evaluation/evaluator.py code served (permissive licence) · get_code("bde3923a8fec6850") |
| score_multi_vector | Ran | illuin-tech/vidore-benchmark/src/vidore_benchmark/evaluation/scoring.py code served (permissive licence) · get_code("1ffacc9d032eaa50") |
| load_vidore_dataset | Not yet run | illuin-tech/vidore-benchmark/src/vidore_benchmark/pipeline_evaluation/dataset_loader.py code served (permissive licence) · get_code("4014d55a46ee32b1") |
Some links come from the archived Papers with Code dataset (CC BY-SA 4.0): attribution and licence.
The ViDoRe Benchmark V1 was approaching saturation with top models exceeding 90% nDCG@5, limiting its ability to discern improvements. ViDoRe Benchmark V2 introduces realistic, challenging retrieval scenarios via blind contextual querying, long and cross-document queries, and a hybrid synthetic and human-in-the-loop query generation process. It comprises four diverse, multilingual datasets and provides clear evaluation instructions. Initial results demonstrate substantial room for advancement and highlight insights on model generalization and multilingual capability. This benchmark is designed as a living resource, inviting community contributions to maintain relevance through future evaluations.
The same record, over MCP at https://syntology.ai/mcp:
get_harvested_code_for_paper("2505.17166")
get_code_for_paper("2505.17166")
have("2505.17166")
Connect an agent — have() is free.