Yambda-5B — A Large-Scale Multi-modal Dataset for Ranking And Retrieval

Published: 5 August 2025| Version 1 | DOI: 10.17632/n2y36w2sv9.1
Contributors:
Alexander Ploshkin, Vladislav Tytskiy, Alexey Pismenny, Vladimir Baikalov, Evgeny Taychinov, Artem Permiakov, Daniil Burlakov, Eugene Krofto, Nikolay Savushkin

Description

The Yambda-5B dataset is a large-scale open database comprising 4.79 billion user-item interactions collected from 1 million users and spanning 9.39 million tracks. This release contains the 50 million‑event subset. Each event is time‑stamped (5 s bins) and labelled with an is_organic flag that separates organic discovery from recommendation‑driven actions, enabling counterfactual and causal studies. The package also includes look‑up tables: album_item_mapping.parquet and artist_item_mapping.parquet – link tracks with their albums and artists. Note: audio‑embedding files are not included in this Figshare release; they remain part of the complete Yambda‑5B dataset on Hugging Face. The full 5 B‑event version (with embeddings) is hosted on Hugging Face. All data are released under the Apache License 2.0.

Files

Institutions

  • Yandex

Categories

Machine Learning, Recommendation System

Licence