A 7 day smart speaker A/B test delivered +4.53% active users and +6.30% total listening time, both statistically significant (p<0.
Yandex Music's old recommender stack read like a relay race. Candidate generators proposed tracks, a pre-ranker trimmed the list, and a ranker ordered what users actually saw. The team's writeup puts the count above fifteen components, fed by large transformer models (Argus, target-attention scorers) and hundreds of hand-engineered features. Sona replaces that whole cascade with a single transformer: one model reads the user's event history, generates candidate tracks, and scores them in the same pass.
The A/B test ran for seven days on Yandex Music smart speakers, with 15% of users in each arm and a production cascade as control. Sona's arm beat the control on two engagement metrics: active users rose 4.53% and total listening time rose 6.30%, both significant at p < 0.01, according to the team's r/MachineLearning post and the Sona Technical Report on arXiv. The team is explicit that this is not yet at full traffic; a longer A/B is still underway.
The mechanism lives in three pieces. First, the encoder reads up to 8,192 user events and splits them into two blocks: an older 6,144-event block and a recent 2,048-event block. Cross-attention and one full-history self-attention layer let the two blocks exchange information before a 7-layer stack runs only on the recent 2,048. The arXiv preprint calls this History Compression and reports it roughly halves inference cost compared with full attention over 8,192 events.
Second, the same encoder output feeds both a decoder, which runs beam search and emits Semantic IDs as candidates, and a Ranking Module that scores those candidates. Semantic IDs are learned discrete codes for each track, so the decoder generates candidates the way a language model generates tokens. The encoder runs once per request. Third, next-token-prediction and distillation objectives update the encoder jointly, so generation and ranking share a user state instead of passing it through a handoff.
Hand-engineered features are gone: the Sona paper states that Sona and its Teacher Ranker both operate on logged event fields and learned item representations, with no hand-built signals. The consolidation is real: one model owns the user state, the candidates, and the order.
Replacing a cascade with one model is not free. The system has to learn the item catalog as part of generation, absorb the ranking signal end-to-end, and train without the safety net of separate models that catch each other's mistakes. The team's bet is that one tightly coupled user state beats fifteen loosely coupled ones on the engagement numbers the A/B measures.
The cost shows up where the old specialization used to live. Catalog coverage, the share of the catalog the system can actually surface, is lower than the production stack, and the team flags this as an unresolved issue they are actively investigating. The cost is openly named: the r/MachineLearning writeup lists it as the open problem. A system that consolidates candidate generation and ranking into one user state inherits a narrower view of long-tail material, because the encoder's pressure to predict the next listen and the ranker's pressure to score well are now the same training signal.
The A/B test does not yet run at full traffic, the result is one platform on one device class, and the numbers are self-reported; no independent audit is visible. The team's longer test, now underway, will report whether the engagement gains hold at full traffic and whether catalog coverage recovers without putting the candidate generators back in.