Back to blog
17 min read

Modern Recommendation Systems

Trends in the last two years of published recommendation systems research

This is the second of two posts on recommendation systems. This one covers modern recommendation systems from a technical perspective. The previous one examined how recommendation system design choices lead to globally significant consequences.

Overlapping title crops from recommendation-systems research papers

For this second post in the series on recommendation systems, I want to focus on how recommendation systems (recsys) have developed in the last two years or so. We’ll walk through a few recently published papers and see what they can tell us about the state of the field. These papers show that ecent recommendation algorithms have taken advantage of the latest in ML, including large AR transformers, deep content encoding, and various efficiency gains. However, they have not converged on a single dominant architecture as we see in language modeling. Instead, companies tailor their systems to their platforms’ specific needs. Ultimately, I want to highlight four trends in recsys research:

  • Modern recsys models are making increasingly heavy use of content features.
  • Many of the field’s largest gains are still coming from scale, motivating extensive research on improving model efficiency and generality.
  • Generative retrieval is growing in popularity, with several major recsys players demonstrating at least partial adoption of the strategy.
  • The recsys field has not consolidated as much as language modeling or image generation. Different companies use disparate strategies to meet their product needs.

Preliminaries

Before getting into the examples, I want to cover two research areas that are important for context. The first is content embedding. Over the last 10 years, there has been a wealth of research on training encoder models for modalities including text, images, videos, and 3D. These encoders have increased in parameter scale, multimodality, and diversity of training tasks, yielding powerful models such as Siglip2, DinoV3, and E5. Taking scale further, we’ve seen LLMs fine-tuned into encoders (e.g. Gemini Embedding and Qwen3 Embedding). For these models, next-token prediction serves as the internet-scale pretraining task, which is followed by task-specific contrastive training to convert the EOS hidden state into a useful embedding. Thanks to all this encoder research, we now have many powerful ways to encode all manner of content.

The evolution of encoders has been driving a slow shift in recsys toward including more and more complex content features. Who engaged with what gives most of the picture, but as we try to milk the last few percentage points of engagement gains, the recsys field has turned toward making the most of the latest and greatest in embedding models for content understanding.

Gemini Embedding 2 maps text, image, video, and audio inputs into a unified vector space
Gemini Embedding
Qwen3 Embedding training pipeline
Qwen3 Embedding training pipeline

The second research area relevant to modern recsys is generative retrieval. Generative retrieval is the idea of treating recommendation like a next-token prediction task over the user engagement sequence. As the entire ML field has moved toward next-token-ifying everything from text to image generation to audio generation, it’s little surprise the same idea is gaining traction for recommendations. While content features, user features, user actions, and other metadata can all be easily represented as vectors or discrete tokens, the hard part lies in tokenizing content. Selecting the next piece of content a user should see as a single token would require an output vocabulary size equal to the number of unique pieces of content on the platform in question. This number is often in the billions, making the pre-softmax output projection impractically large. A common way to solve this is multi-hashing: hash each piece of content kk times into NN buckets, yielding a probably-unique kk-tuple of integers less than NN. The decoder-only GR model can then output kk tokens for each piece of content, each with a manageable vocabulary size.

We can sometimes do a bit better by using semantic IDs instead of mult-ihashing. Like multi-hashes, semantic IDs are learned mappings from the set of all content items to kk-tuples, but they enforce a semantic hierarchy on the resulting tokens. That is to say, if two pieces of content share the same first token, they should be semantically similar. If they share the same first two tokens, they should be even more similar. Semantic IDs work by training an RQVAE on top of content embeddings. The RQVAE’s code indices form the kk-tuple used to represent each item.

Generative retrieval on semantic IDs was popularized by the TIGER paper, which got a lot of buzz after its publication. However, multiple later works tested TIGER-like systems in different settings and found them lacking. Issues included poor cold-start handling, pathological semantic ID hierarchies, lost representational capability, expensive inference, and more. Several of these works propose systems that keep the sequence modeling task of GR, but maintain an ANN-search for selecting the next piece of content to recommend. This is unsurprising for two reasons. Firstly, ANN-search-based methods are uniquely well-suited to the problems of retrieval, while next-token prediction is far from an obvious fit for the task.

The second reason is a hotter take. Having spent much of the last three years at Roblox working on audio and translation ML, I’m familiar with RQVAE and transformer decoding—two critical components of GR and semantic IDs. More importantly, I’m familiar with what a pain it is to make both of these work. The RQVAE audio codec has been reinvented by reputable researchers dozens of times since its near-simultaneous invention by Google and Meta in 2022. No one has managed to agree on which codec is best, what strategy produces a good codec, or even if we should use RQVAE codecs at all. Making RQVAE converge is miserable. So is making their latent space “nice” in all the ways we might want it to be for a task like GR. Similarly, decoding can quickly become a nightmare. From minimum Bayes risk to beam search to top-k to top-p to min-p sampling, there are so many ways to get what you’re looking for out of an AR transformer, and usually only one or two of them work well. Due to the engineering complexity of AR decoding and semantic IDs, I expect these methods may become more popular in several years once researchers have had the time to develop the black magic implementation details required to get them to work. We’ve seen similar narrative arcs for other finicky model types like flow matching, decoder-only transformers, and discrete diffusion. I may be completely wrong here, but I wouldn’t count pure GR out entirely just yet.

TIGER generative retrieval with semantic IDs
TIGER generative retrieval
Residual quantized variational autoencoder producing semantic codes
RQVAE semantic IDs

Case Study 1: Pinterest

Pinterest publishes more recsys research than almost any other company and has developed a strong reputation in the field. Their papers tend to give more detail than competitors, which those of us outside big recsys companies may appreciate.

Pinterest first published multiple works purely on content encoding, culminating in “OmniSage: Large Scale, Multi-Entity Heterogeneous Graph Representation Learning,” their model for embedding pins, users, engagement sequences, and anything else on their platform. They create a graph of all relevant entities (the paper avoids being explicit about what these are beyond users, pins, and boards) and embed each node along with context from nodes in its neighborhood. They train on several contrastive tasks, leveraging several system-level optimizations and one backprop optimization to be able to handle their very large graph size. In A/B tests, they observe several statistically significant engagement gains from employing OmniSage’s learned representations for ranking and retrieval in production. This isn’t a novel research breakthrough, but it does show that Pinterest put a lot of effort into high-quality content embedding, and, more importantly, better content understanding can move engagement metrics.

OmniSage architecture with transformer decoder and multiple contrastive-learning tasks
OmniSage architecture

Pinterest followed OmniSage with “PinRec: Unified Generative Retrieval for Pinterest Recommender Systems,” a retrieval model that hybridizes generative retrieval and ANN search. They take a user’s sequence of engagements up to time tt, embed them with OmniSage, and shove the whole thing through a transformer. They then use a contrastive learning task to make the last timestep’s hidden state close to the OmniSage embeddings of timesteps tt through t+nt+n, with n>1n>1 allowing the targets to occur in slightly different orders. They fine-tune separate versions of PinRec for the Search, Related, and Home Feed surfaces. At inference time, PinRec does an ANN search over pins’ OmniSage embeddings using the transformer’s last hidden state as a query. It then plugs the selected neighbor in autoregressively as the input for the next timestep. This differs from typical autoregressive decoding where the network predicts a discrete token that is passed through an embedding layer to obtain the input for the next timestep. Instead, the network produces a continuous embedding, which is discretized to select a single piece of content via ANN search.

The PinRec authors claim in the paper that they tried using semantic IDs as their input representation, but found that they suffered from “representational collapse” and opted to use dense vectors instead. Especially for input representations, this is unsurprising. On the input side of a transformer, you’re able to use dense vectors, and any representation derived from them incorporating no new information is unlikely to be more powerful. Without semantic IDs or multi-hashing, they need some way to turn a transformer’s continuous output into a choice of a piece of content, which is where the ANN search comes into play. Combining AR transformer decoders with ANN search for selection makes use of the last 8 years of work on scaling transformers without flipping on its head the recsys formulation of a company heavily reliant on recommendations.

PinRec processes Pinterest engagement sequences to retrieve recommended pins
PinRec

The final Pinterest paper in this sequence is “UniPinRec: Unifying Generative Retrieval and Ranking at Pinterest Scale.” In UniPinRec, the authors augment PinRec to make it into both a retriever and ranker in a single model. UniPinRec changes several training details from PinRec, but there is one very simple core improvement: They add an extra MLP head that predicts the probability of each user action on the piece of content provided as input to the transformer at time tt. That’s it. This motivates about a page in the paper on the new training strategy necessary to learn both next-embedding prediction and action prediction, but the entire rest of the paper is just about serving optimization and evaluation. That’s what this paper is really about: making recsys more efficient. The huge cost-saver here is that user action prediction is a ranking task; reweight the action probabilities and you have a score to rank with. This means ranking just involves one more forward pass of the retrieval model after the model completes PinRec-style autoregressive retrieval. The K/V cache from retrieval can be re-used, making ranking extremely cheap.

UniPinRec unifies recommendation-system input formats, architecture, training, and serving
UniPinRec

Pinterest’s recent work made gains in content understanding and scale, while mixing GR’s AR prediction with existing ANN retrieval. We can also see how the interchangeability of content on Pinterest affects their design. The only place where the recommendations stack has a chance to memorize a specific piece of content is at the OmniSage training layer. Beyond that, there is no use of hashing, semantic IDs, or memorized embeddings to enable UniPinRec to recommend a specific piece of content from its corpus. Instead, UniPinRec foregrounds deep content and engagement understanding to recommend any pieces of content that are most relevant to a user, illustrating how a platform’s content landscape shapes its recsys architecture.

Case Study 2: Meta

Next, we’ll take a look at two papers from Meta that point in a similar direction to Pinterest’s work. The first, “Unifying Generative and Dense Retrieval for Sequential Recommendation,” is a direct response to TIGER, showing lackluster results for the TIGER architecture on several datasets, especially for cold start. They propose two simple modifications, naming their new system LIGER. First, they add dense embeddings to the semantic IDs used to represent content as model inputs. Second, they use standard ANN retrieval specifically for cold start content. They find this lifts offline metrics on several research datasets. It’s important to qualify that both TIGER and LIGER are research systems proven exclusively on offline results. The papers exist to develop ideas. They do not claim their described systems perform best in the wild. Regardless, this Meta paper demonstrates the trend of pushing back on pure GR while still keeping some of its core methods.

LIGER combines dense and generative retrieval with semantic ID generation
LIGER

Meta’s more relevant paper, “Trillion-Parameter Sequential Transducers for Generative Recommendations,” introduces a production-scale model they call Hierarchical Sequential Transduction Units, or HSTU. HSTU is very similar to UniPinRec: It pushes a long sequence of user-engaged pieces of content and actions through a large transformer-like architecture, emits continuous embeddings, performs an ANN search over content embeddings to select a piece of content for retrieval, and then autoregresses. However, rather than predicting actions with a separate head at each timestep, they simply interleave actions with content in the predicted sequence. The paper makes a big deal of its novel HSTU architecture, which turns out to be a transformer with a linear layer and the attention softmax removed. The removal of the softmax normalization on attention scores allows the network to do a better job of counting its inputs’ sequence length, which is important for retrieval tasks. Their more significant contributions are a slew of systems optimizations that exploit sparsity, reduce memory use, and improve parallelism. The last major contribution of this work is an analysis of scaling laws. They test parameter scales up to 1.5 trillion, an experiment only a company of Meta’s size could dream of paying for, showing that for their GR-style retriever and ranker, metrics improve linearly with the log of compute. Significantly, they find this scaling does not hold for traditional retrieval methods. The paper includes online A/B tests, demonstrating significant metrics wins.

HSTU highlights the same convergence as UniPinRec toward single-transformer merged ranker and retriever architectures. They both keep ANN search as the core retrieval mechanism, but adopt GR-style autoregression. These more general architectures benefit from increased parameter and data scale, which becomes more affordable every year. Efficiency becomes the most important axis of further improvements, since it enables scale, which in turn yields predictable performance gains.

HSTU generative recommender scaling laws compared with traditional deep learning recommendation models
HSTU scaling laws

Case Study 3: Spotify

Now, we’ll look at one example that bucks the trend: “Deploying Semantic ID-based Generative Retrieval for Large-Scale Podcast Discovery at Spotify,” which introduces a system called “GLIDE.” This is the only example I’ve seen of a large recsys-based company wholesale adopting semantic IDs and GR. Granted, this model is exclusively for podcast recommendation, which is a narrow surface within Spotify’s broader product. It’s still interesting to see how and why Spotify would go for an unpopular technique here.

Their model is fine-tuned from an off-the-shelf Llama 3.2-1B. They first fit a semantic ID RQVAE on podcast names and descriptions, then augment the transformer’s vocabulary with the semantic ID tokens. They pretrain the semantic ID embedding layer using a translation objective, mapping podcast semantic IDs to and from textual descriptions of the podcasts. After that, they add LoRA adapters and train on the actual recommendation task. To obtain different flavors of recommendations, they prompt the GR model with simple text prompts. The possibility of being able to guide a retrieval model with arbitrary text prompts is exciting, but in practice GLIDE only trains on a few simple prompt templates, undercutting the utility of prompting. It would be interesting to see future work that builds large-scale synthetic datasets for promptable recommendation to support real text-guided recsys.

Several implementation details reveal the difficulties Spotify encountered with generative retrieval and semantic IDs. They resample training sequences to suppress the most popular pieces of content and upweight training sequences exhibiting more exploration-heavy behavior. They also break semantic ID ties (situations where two pieces of content have the same semantic ID) using popularity, meaning that for pieces of content sharing the same semantic IDs, only one of them will ever be recommended. Spotify ultimately tested GLIDE in production and reported statistically significant metric wins.

GLIDE takes a meaningful step toward productionizing a true semantic ID + GR framework, but remains a long way from a foundational recs model like UniPinRec or HSTU. Text-prompted recommendations are still more wishful thinking than reality, and the weak tie-breaking mechanism is a serious flaw. Using only textual descriptions of podcasts as content features is also limiting. That said, the GR + semantic ID framework does appear to be a good fit for the podcast recommendation use case. Podcasts and songs are not interchangeable even when they have similar content, making memorization ability critical. Cold start is also less important than other domains, since podcasts are slow to produce and run for multiple years. The ability to reason in natural language and support multiple use cases via natural language prompting could also eventually be useful for multi-surface adoption. Spotify exhibits recommendations on a variety of surfaces and sorts on its pages, and often uses generated titles for playlists. With some further work, one model may be able to support recommendations, playlist generation, and playlist titling for the whole platform.

Case Study 4: X (Twitter)

Lastly, we take a look at the open source version of the X recommendation algorithm. We should be wary of the fact that the open source version may be far from what’s operating in production, but as we have no other material to go off of, the repo will have to do.

X’s algorithm appears in many ways surprisingly behind the times. It does not use content features whatsoever, simply embedding users, authors, and posts with a multi-hashing scheme. The retrieval model is a standard two-tower architecture with a transformer that ingests the user engagement sequence on the user side. The ranker uses the same transformer backbone to encode user features and then predicts the probability of various engagement actions. A manual weighting of the action probabilities produces a final ranking score.

The omission of content features indicates a very strong inductive bias toward memorization. The system is able to easily memorize specific tweets to recommend, but abandons interchangeability of tweets with similar content. The lack of content features also makes cold start recs intractable. For X’s case this may be salvageable given every tweet is first engaged with by followers of the author, providing a useful learning signal at the start of a tweet’s lifecycle. Regardless, the simple approach is surprising for a company that relies on recsys as heavily as X does. Their previous open source release in 2023 demonstrated a much more complex, content-heavy algorithm. The 2026 version at the very least illustrates a move toward more general architectures that likely benefit from the broad efficiency gains in transformers developed over the last several years.

Researchers, journalists, and users have reported several changes in X’s recommendations, including a rightward political skew and a greater concentration of engagement among a smaller set of accounts. This could be an unforeseen outcome of building what from the public repo seems to be a subpar recommendation system. However, given Musk’s acquisition of X and subsequent efforts to use his immense wealth to realize his political goals with disastrous consequences, I would leave less friendly interpretations on the table. The open source implementation may be slop produced as a PR stunt to perform transparency while bearing no resemblance to what’s running in production. The production system may also contain explicit rules to bias it toward its owner’s agenda that are not reflected in the repo. The main point of my first recsys blog post was that most of the time, harmful consequences of recsys are not the developers’ direct intent. This is one of the rare cases where I would suspect the opposite. At the same time, we cannot know with certainty whether the observed bias is intentional or incidental, and, in the end, I don’t think it matters much. In either case, it has deleterious effects on X’s community and individual users.

Conclusion

If you made it this far, thanks! This one was pretty dry. The most important takeaways for the state of modern recsys come from the Pinterest and Meta examples. Large recsys companies are moving toward generic, single-transformer architectures that benefit from friendly scaling laws and a plethora of inference optimizations. These models adopt autoregressive inference, the most important aspect of generative retrieval. They make heavy use of content features, leveraging large encoders to embed these platforms’ multimodal user-generated content. At the same time, the field is far from having landed on a single dominant architecture. Since different platforms are home to very different types of content and exhibit varied social dynamics, I believe the field is likely to remain diverse even as the rest of machine learning has increasingly consolidated around a smaller set of architectures. It’s always exciting to see papers like GLIDE that make an unpopular method work in production.

While I do find recsys full of technical intrigue, this does not diminish all the social consequences discussed in the previous post. As recsys models grow bigger and more general, they become even less interpretable. The likelihood that undetected pathological behaviors will arise continues to grow. Developing methods for the evaluation of recsys safety and ethics remains as important a task as ever. Furthermore, the X algorithm case demonstrates that while it is far from the norm, there is still room for company execs to explicitly bias their organizations’ algorithms. But regardless of the source of recsys’ undesirable behaviors, their impact on our lives at both individual and societal scales remains the same. Although we can never be immune to their seductions, I hope that understanding of how recommendation systems work can help us all become a bit less prone to manipulation by The Algorithm.

References

Acknowledgements

I used ChatGPT and Claude for feedback and edits.

Discussion

OMG there's even a chat, this is basically Reddit