Yandex

Yandex Open-Sources YTsaurus Flow for Real-Time Data Processing

October 06, 2026

Yandex has open-sourced YTsaurus Flow, a technology for continuously processing large data streams in real time. Already used across Yandex services, it processes more than 100 GB of data per second in Yandex Ads alone, including millions of online events such as clicks, impressions, likes, and orders.

By processing these signals as they occur, YTsaurus Flow helps systems respond more quickly to changes in user interests, detect malicious activity, improve recommendations, and select more relevant advertising.

Why real-time data matters

User interests and needs can change quickly. Someone might search for a flight and then move on to a new movie or products for their home. To recognize these shifts, algorithms need fresh data about user activity — what people buy, what music they listen to, or what they search for.

Traditionally, this data is often accumulated and processed in batches rather than immediately, which can introduce delays of several hours. YTsaurus Flow reduces this delay to nearly zero by processing data continuously as it arrives.

The technology can be used in a wide range of applications where data freshness matters, beyond advertising and recommendation systems. For example, it can prepare data for training machine learning models, update information on websites and in applications, and track views, clicks, and orders.

YTsaurus Flow is designed for high-load environments and provides exactly-once processing, ensuring that events are neither lost nor duplicated. It also continues operating in the event of failures by redistributing workloads across available servers.

How Yandex uses YTsaurus Flow

YTsaurus Flow is already used across a number of Yandex services. At Yandex Market, it helps collect statistics faster so sellers can respond more quickly to changes in demand. Yandex antifraud systems use it to detect malicious activity, while Yandex Ads uses it to select more relevant advertising.

Previously, preparing data to further train some models used for ad selection could take more than 10 hours. With YTsaurus Flow, this delay has been almost entirely eliminated. Within Yandex Ads alone, the technology processes more than 100 GB of data every second, including millions of online events such as clicks, impressions, likes, and orders.

YTsaurus Flow operates as part of YTsaurus, Yandex's open-source platform for storing and processing large volumes of data, designed for large businesses. With YTsaurus Flow, companies can process data both in batches and continuously as it arrives. They can also work with historical and fresh data within the same system.

YTsaurus Flow was jointly developed by the Yandex Infrastructure and Yandex Ads teams. The technology is available under the Apache 2.0 license, which allows it to be used in commercial projects.

IPJSC “Yandex”

Head office
16, Leo Tolstoy St., Moscow, Russia 119021
Investor Relations
Public Relations
Corporate Secretary