The Banality of Evil, or Why Your Feed is so Fucked Up
How ordinary engineering decisions create the recommendation engines we all know and hate
This is the first of two posts on recommendation systems. This one examines how recommendation system design choices lead to globally significant consequences. The second will cover modern recommendation systems from a more technical perspective.

AI models that generate text have been busy reshaping our world since the release of ChatGPT in late 2022. In this post, I want to talk about another class of AI models that have also had a profound impact on all of our lives: recommendation systems (recsys for short). Many of the biggest tech companies, including Google, Meta, Netflix, and ByteDance, derive the majority of their revenue from either recsys-based products or personalized advertisements, which at their core are just another type of content recommendation. Even search is highly adjacent: recommend the best page, product, etc. given user data and a query.
Over the past two decades, recsys algorithms have shaped political revolutions, become a battleground of many elections, ruined our attention spans, and built some of the world’s most powerful companies. On a more personal level, these algorithms recommend content that informs the way we speak, the pastimes we pick, the way we date, what we find funny, and practically every other belief we hold. If not for the YouTube recs algorithm, I would be a vastly worse guitarist, know nothing of periodization for strength training, and have spared myself hundreds of hours now irretrievably sunk into spectating the chaotic world of competitive armwrestling. Recommendation systems understand some aspects of us frighteningly well, sometimes better than we do ourselves. They can predict what we are likely to talk about with such accuracy that many swear their phones must be listening to them.
Despite their influence, recommender systems are poorly understood outside of the circles of people that build them. News tells us that social media feeds skew toward inflammatory, divisive content, prompting ire at tech companies. For many, this may conjure up images of villainous executives sitting around a table theorizing how sowing division will make them more money. The image gives us an easy scapegoat, but is likely far from the truth. Problematic outcomes of recsys in the wild tend to be side-effects of much more mundane engineering decisions.


From Implementation to Impact
Objectives
The goal of any recommendation system is to match users with pieces of content. More formally, if we have users and pieces of content (both and sometimes being in the billions), we want to pick pieces of content (, say 10 or 100) that maximize some objective. The most important decision in building any recommender system is the choice of that objective. Many platforms choose to maximize some notion of “engagement.” But what is engagement? Clicks? Likes? Follows? Minutes of activity? ? There is no obvious answer. And yet the choice of one over another will produce a remarkably different recommendation landscape.
When I worked on search systems at Roblox, the discovery team had chosen to optimize for two targets: “engagement” and “relevance.” I won’t share the exact definition they had for engagement, but I assure you it would sound as arbitrary as any of the examples I gave above. A reasonable question might be: If the goal for the company is to make money via engagement, why would we optimize for “relevance,” whatever that is?
Taking a step back, the real goal of any corporation is something along the lines of making as much money as possible over the next several decades. The estimation of this objective is completely intractable, so we use proxies for it. A more obvious chain of proxies might be something like:
make money in 20 years → make money in 1 week → maximize time spent on the platform in 1 week
Weekly timespend is much easier to measure and train models against than profit 20 years from now. But what we found at Roblox was that search algorithms trained to maximize near-term engagement tended to boost the most popular games at the time along with a whole lot of clickbait. We couldn’t be certain, but we hypothesized that this kind of behavior would disincentivize the creation of new high quality content, hurt long-term retention, and ultimately damage the Roblox ecosystem. In other words, we thought there was a meaningful gap between the real objective “profit in 20 years” and the proxy “engagement now”. To fill this gap, another proxy was born: “relevance”.
The idea that surfacing more relevant content in search would help was not backed by all that much data, but even to an uninformed observer, it would seem pretty reasonable. The team sketched out a list of criteria that defined what it meant for a game to be relevant to a search query and then had teams of paid human labelers rate the relevance of tens of thousands of game-query pairs. We trained models on this data that learned to estimate relevance, which were then integrated into Roblox’s search algorithm. I’d like to note the whole methodology is not novel and has been replicated in various forms at other companies (e.g. Pinterest). Look at the long chain of proxies here:
profit in 20 years → relevance → what Roblox players would find relevant → what paid human labelers following our ruleset thought was relevant → a model estimating the human labels
The model we get is the result of a long chain of rough approximations, each of which has an opaque but significant influence on the nature of Roblox search.


The key takeaway from the Roblox example is that the objective a recommendation system is trained against is a choice made by engineers. It is not a fact that falls out logically from the nature of the company or platform, nor is it a mysterious principle defined by evil-minded executives. It is the mundane result of a deep stack of inferences, assumptions, and proxies guided by imperfect data and decided by teams of engineers and their managers. When we find that our favorite social media site is biased toward inflammatory posts or a certain political ideology, that usually wasn’t an intentional decision on the part of the company. More likely, it is something a complex model learned to do after training on a lot of data to maximize some objective like “sessions with 5+ minute video watch times” that was decided upon because messy trails of evidence led the team to believe that was a good proxy for profit years in the future.
Technical Challenges
Once we have an objective, the next problem becomes training a model to maximize it. The data we have to work with includes user metadata, the content itself, the sequence of contents a user has previously engaged with, and any other general domain internet data. The datasets involved can be enormous, and they are constantly changing: users join and leave platforms, content is created and deleted, and over long time horizons, user preferences start to shift. This warrants regular—often daily—retraining of recsys models. AI models are usually very expensive to train (a single training run of a large LLM costs millions of dollars), but due to retraining requirements, recsys models cannot afford the same luxuries.
Inference time, when the model is used to generate live recommendations for users, brings another slew of challenges. The number of unique pieces of content available to recommend, or our “vocabulary size,” often figures in the billions. It is hard to decide what to do with new content boasting little past user interaction data to train on (the “cold start” problem). The model also needs to be able to handle millions or billions of requests per day without breaking the bank on compute costs.
Deeper problems arise when we try to ask if our model is serving up high quality recommendations that benefit our platform’s ecosystem. Since recsys models are trained on user interaction data that occurred on top of recommendations generated by previous versions of the model, all kinds of degenerate self-reinforcing and drifting behaviors can arise over long periods of time. A whole science of A/B testing and experimentation has emerged, built around the measurement of recsys performance. Impacts of a new algorithm are nuanced and metric gains are often very small, motivating decades of data science and statistics research on understanding how recsys models are affecting user behavior.
All this goes to say that even beyond deciding on an objective, there lurk a wealth of complex design choices with multifaceted consequences that could yield some of the pathological recsys behaviors we see in the wild. A tremendous amount of money and decades of cumulative research effort has gone toward deciding what short video you should watch next in order to maximally profit off you. In light of all this machinery, it should not be surprising when you get an ad for something you were just talking about, nor should it be surprising when TikTok’s recs algorithm autonomously discovers it can farm you for more engagement by weaponizing your insecurities and turning you into an incel. It may not succeed, but hell, it’s certainly worth a shot.

Recsys Crash Course
In this section, I want to cover more technical information on how recommendation systems work that may provide some color for the earlier arguments as well as some setup for the second blog post. This will not be a comprehensive introduction to recsys, nor will it be an attempt at providing high-level intuitions. To address the latter, my former colleague wrote an excellent post here I’d strongly recommend checking out before reading further. Below, I will try to connect the math and diagrams to engineering design choices and outcomes discussed above.
Retrieval and Ranking
Most recommendation systems contain two distinct stages: retrieval and ranking. The first stage, retrieval, is where the system, given a user, retrieves a large corpus of candidate items that it estimates to be relevant to the user. Retrieval is what controls the universe of content the platform deems might be worth some of your time. During ranking, the system assigns scores to each candidate, allowing them to be ranked in order of relevance to the user. Ranking is where the algorithm gets an extremely powerful say over what you’re likely to consume by deciding what comes first. It is also where small implementation differences can yield the greatest changes in user experience.
Retrieval models often operate on vector similarity. They encode each user and piece of content as a vector in a high-dimensional space. Given a user, they take its associated vector and find the pieces of content with vectors nearest to the user vector. This nearest-neighbor search becomes computationally expensive for large numbers of users and contents, so there is a broad class of algorithms under the category of Approximate Nearest Neighbor Search (ANN Search) that make it faster. The setup is convenient because it allows fast retrieval of relevant content without having to re-encode each piece of content as a vector every time we want to make a recommendation.
Since retrieval doesn’t have to abide by any ordering, multiple retrieval algorithms can be ensembled together. We call each of these retrieval mechanisms a Candidate Generator (CG). While some CGs, like ANN-based ones described earlier, are more complex, others can be very simple. For example, one CG could return posts from people your account follows on a social media platform, or posts that have a keyword appearing in your bio. Adding candidate generators with different focus areas is one way recsys creators can bias recommendations toward certain types of content. For instance, a game platform CG returning games with thumbnails that are semantically similar to a user’s bio text may favor games with busy thumbnail images, as they are likely to match with a wider variety of texts. We typically evaluate CGs using metrics like recall, which roughly counts, “out of all the content the user engaged with, what percentage were in the corpus returned by the CG?” At the same time, there is an incentive to limit the number of generated candidates to make the ranker’s job easier.

Ranking models whittle down the corpus returned by the retrieval stage into a ranked feed like you might see on most social media homepage sites. A ranking model will typically take as input a user and a single piece of content, and then output probabilities of a set of predefined user actions occurring. For example, given Nameer Hirschkind and a video of armwrestling champion Levan Saginashvili, a ranker might estimate:
A score is then derived from these probabilities with a weighted sum:
The choice of those weights is one of the many key choices of objective a ranking algorithm’s creators must make. This choice directly reflects the priorities of the company behind the algorithm, and drastically affects what kind of content users see. A really simple example of this is how increasing the weight on would likely boost clickbait content that users do not engage with deeply.
Since ranking models only have to consider the small pool of candidates returned by retrieval rather than the whole content corpus, they can afford to do more computation per sample and make more nuanced distinctions. We typically evaluate ranking models using more complicated metrics that compare the model’s ranked sequence to the sequence of items a user chose to interact with or ignore (examples are NDCG, DCG, mAP).

Common Methods
To close the loop, we’ll go over a couple concrete examples of recsys algorithms for retrieval and ranking. While the details of ranking algorithms can get very deep, ranking fundamentals are a bit more straightforward, so we’ll spend more time on retrieval.
The first and most basic retrieval method we’ll cover highlights an important point in recsys: Information about the actual content users are consuming, such as its text, image, or video makeup, is secondary. Who engaged with a piece of content tends to be much more informative than the content itself. As a human, this point is deeply unintuitive. We tend to base our decisions of what content to engage with almost entirely on the content itself. Recommendation systems are very different.
The method in question is Matrix Factorization (MF). Supposing we have users and pieces of content, we set up an matrix whose entries record whether a user engaged with a piece of content:
We then pick an integer that we will use for the dimension of the user and content embeddings (an embedding is a semantically meaningful vector representation of something). We approximately factorize into the product , where is tall and thin and is short and wide:
Each row of gives us a user embedding while each column of gives us a content embedding. We can also scale these vectors during the factorization so they are all of length . When you dot product two vectors of length , the resulting number is the cosine of the angle between them:
This is only when that angle is , and it moves toward as that angle grows to . We have set up our factorization so the angle between a user embedding and content embedding should be near when a user is likely to engage with that content. Given a user embedding, we can then perform ANN search over all the content embeddings using the dot product (i.e. cosine of the angle between the vectors) as our distance metric to obtain content recommendations.

Matrix factorization is an old, simple, and naive method for retrieval. It does not take into account anything about a user or piece of content. It only considers who engaged with what. The numerical methods behind approximate MF have existed for over 100 years. And yet, in practice, MF works pretty well. A 2026 Netflix paper estimated that swapping their recs algorithm for random recommendations would decrease engagement by 16% while switching to MF would only decrease engagement by 4%. The thing to remember here is that “who engaged with what” is extremely powerful information at scale, even when paired with the most basic machinery.
A more modern framework for retrieval is the “two tower” class of models, which are still widely used in the industry today. Two tower models enable a significant upgrade over MF: they can understand content. They have the same underlying objective: learn high quality embeddings for users and pieces of content, enabling ANN search for retrieval at inference time. They involve two neural networks: a “user tower” and a “content tower”. Each tower takes as input some sort of unique ID for the user or piece of content, along with any other information the model’s creator thinks is important. For the user tower, this could be demographic data, content engaged within the last 7 days, or cumulative platform engagement statistics. For the content tower this could be content metadata (date of upload, size, etc.) or elements of the content itself, such as its text, images, video, or whatever else comprises it. Each tower outputs an embedding vector. In the vanilla version of the model, the networks are then trained to optimize a simple objective: when a user has engaged with a piece of content, make the dot product of the embeddings close to 1. There are several formulations of this objective, including triplet loss, InfoNCE, and plain old binary crossentropy, but all of them support the same general task.

Two tower models have remained popular because of their versatility. The user and content tower architectures can become complicated, taking advantage of all the latest developments in AI. The model’s developer is free to modify the objective function to push user and content vectors together or apart under more nuanced sets of circumstances. In cases where the number of contents is small, the model can memorize details about each piece of content, but it is also able to generalize given rich content features such as text or images. Ultimately, the more flexible framework enables far more options for recsys developers. Choices of objective, network architectures, and what inputs to use for both users and content will contribute to the behavior of the system.
Ranking is simpler to formulate. Ranking models receive as input any information about a user and a piece of content, and return a score. Remember that information about a user can (and usually should) include some kind of engagement sequence. A common strategy is to use some kind of neural network to take in the aforementioned inputs and produce estimated probabilities that a user will perform different modes of engagement with a piece of content (e.g. click, like, repost). Engineers can then manually weight those scores according to the company’s goals and treat the sum as the final ranking score. We have discussed how this weighting has an outsized effect on recommendation results. Contents are surfaced to users in order of their ranking scores. Ranking models can scale in complexity by making use of more diverse information about the user and contents or via adding complexity to their neural network architectures.
Conclusion
We’re done with the boring part! Details aside, I hope this section gave some examples that concretized how specific design decisions in retrieval and ranking might sway recsys behavior. I also want to reiterate the point “who engaged with something tells you more about it than the thing itself.” It’s unintuitive, but core to how recommendation systems behave. Next time you’re scrolling through a feed of recommended content, as we are all prone to doing, it’s worth remembering the decades of research, billions of dollars, and heaps of machinery carefully curating each selection. The placement of every degenerate reel is the product of a highly engineered system, and yet the reasoning behind it is largely opaque, lost in the weights and activations of large models. The choices that led to the creation of such a system were intentional, but mostly orthogonal to the personal or political downstream impacts.
One message I don’t want to send is that the makers of recommendation systems should be absolved of their creations’ crimes. Rather, it is that the ill effects of recsys-based products are not to be blamed on a few bad actors. They are the product of the technology and the entire incentive structure they are built in. Striving for the goal of “show people videos they like” is more than enough to do irreparable damage all over the world. In the Roblox case, we were fortunate: maximizing ecosystem health via relevance seemed to align with increasing shareholder value. But we aren’t usually this lucky. In many cases, optimizing for what’s best for the company will produce recommendation systems that sow division, prey on insecurities, or propagate dangerous ideologies. What's best for the world and what's best for the bottom line often genuinely diverge. When monetary incentives don’t encourage building responsible recsys, there aren’t many places to turn. I have precious little faith in the ability of governments to regulate technically deep systems. As individuals we have some power. We can choose where we spend our engagement while educating ourselves enough to be aware of the algorithmic pulls on our minds. Unfortunately, that sounds exhausting. Perhaps I’ll feel rested and ready after vegetating with some reels.
References
- Ye et al., “Auditing Political Exposure Bias: Algorithmic Amplification on Twitter During the 2024 U.S. Presidential Election” (2025)
- Wang et al., “LLM-based Relevance Assessment for Web-Scale Search Evaluation at Pinterest” (2025)
- Chaney, Stewart, and Engelhardt, “How Algorithmic Confounding in Recommendation Systems Increases Homogeneity and Decreases Utility” (2018)
- Gilotte et al., “Offline A/B testing for Recommender Systems” (2018)
- Koren, Bell, and Volinsky, “Matrix Factorization Techniques for Recommender Systems” (2009)
- Zielnicki et al., “The Value of Personalized Recommendations: Evidence from Netflix” (2025)
- van den Oord, Li, and Vinyals, “Representation Learning with Contrastive Predictive Coding” (2018)
Acknowledgements
I used ChatGPT and Claude for feedback and edits. Huge thanks to my former colleagues at Roblox; you know who you are. I'm so grateful to have been able to learn from and work with you.
Discussion
OMG there's even a chat, this is basically Reddit