During my freshman year, my school held an assembly with a former TikTok employee who had worked on its recommendation algorithm. He was meant to talk about how social media is bad for us and why we should delete it; he even successfully convinced some of my friends. Unfortunately, I accidentally fell asleep and missed the entire spiel, so I did not delete my one social-media app, Instagram.

I continue to use Instagram, more than I’d like to admit, and do not currently have plans on stopping. I know that its feed is designed to keep me watching. Knowing this, however, has not made it any less effective.

So what exactly makes the feed so effective at monopolizing your attention? Simply put, it’s trained to be effective. That may seem like a circular conclusion, but that’s truly what the entire recommendation system boils down to.

Image illustrating recommendation systems
The author’s Instagram Explore page, where recommended posts are selected from a much larger pool of content and ranked according to what the system predicts the user may want to see. Screenshot by the author.

How the Feed Narrows the Field

Social-media platforms use recommendation systems to select and order the content most likely to interest each user. Most large-scale social-media and streaming platforms (which use similar types of recommender systems) employ deep-learning models to curate your personalized feed. Deep learning is a subset of machine learning that uses neural networks with multiple processing layers to learn increasingly complex patterns and representations from large amounts of data.[1] These networks were loosely inspired by the structure of the brain, although they are not realistic models of biological neurons.[2] Most large recommendation systems divide the process into two main stages: candidate generation, also called retrieval, and ranking. This lets the system quickly narrow millions of possible items to a manageable group before applying more detailed models. In the candidate generation or retrieval stage, the system narrows millions of available content items to hundreds or thousands of potentially relevant candidates. The ranking stage takes these candidates and refines the list to the items with the highest predicted relevance/engagement by giving them a score reflecting platform-specific objectives before returning the highest-ranking set, the final number varying by platform.[3]

Instagram, specifically its Explore page, uses this general framework with a two-tower neural network implementation. In its retrieval stage, the deep-learning model is split into two “towers” or separate models. One tower is dedicated to the user and takes user features as input data. In recommendation systems, this generally includes past interactions, followed accounts, recent activity, and other available user signals. The model compresses this input into a fixed-length vector embedding representative of the model’s current estimate of the user’s interests or preferences at that moment. The other tower does the same for content items, such as photo and video posts. It takes item features as input, which in recommendation systems can include information about the post, creator, topic, format, and engagement history. It then creates a compatible fixed-length embedding representing the item rather than the user.[4]

In order to estimate whether a user will engage in some capacity with an item, a similarity score is calculated between the user and item embeddings. In recommendation systems this is often done by applying the dot product.[5] The dot product of user and item embeddings would by calculated by multiplying their corresponding values and adding the products. Because both embeddings have the same number of dimensions, each value in one vector can be paired with a value in the other. A higher score suggests a stronger predicted match between the user and the item.

As part of the retrieval process, for a given user’s feed, the system could calculate the dot product between their embedding and every available item embedding. However, the use of the two-tower implementation allows this process to be sped up through precomputation and caching. Because user and item features are not mixed as part of the input, item embeddings can be generated once per day and cached for use across many users. Furthermore, these precomputed item embeddings can be stored in a service that supports online approximate nearest-neighbor search. Baking videos, for example, may have embeddings located near one another because the model has learned similar patterns from them. The service organizes the stored embeddings so it can quickly search for those closest to a user’s embedding. This avoids comparing the user embedding individually with every available item. Since nearby embeddings represent stronger predicted matches, the nearest items can then be passed on as candidates for ranking. To continue with the baking example, if a user frequently engages with baking videos, this will be numerically reflected in their embedding so that it can be matched with nearby baking-video embeddings. The similarity comes from the overall pattern of values in the embeddings, rather than one specific value directly representing baking.

Instagram can also retrieve candidates by finding items similar to posts a user has previously liked, saved, or shared, rather than relying only on the user embedding. This is possible because, with the two-tower implementation, the item features are processed independently, so the system can find candidates using only item embeddings. Generally, recommendation systems have more than one way of filtering down candidates in the retrieval phase.[6]

Illustration of recommendation ranking
A simplified view of retrieval in embedding space. The system represents users and content as numerical embeddings, then retrieves nearby content embeddings as candidate recommendations for ranking. Illustration created by the author with assistance from ChatGPT.

From Candidates to a Final Ranking

With the Explore page, the ranking stage is also split in two, since even after the retrieval stage there are too many candidates to use a heavy model on all of them. Heavy models can make more precise predictions and account for more complex relationships between features, but require far more computing power and are therefore slower. As a result, both a first-stage ranker, which is a lightweight model, and a second-stage ranker, which is a heavy model, are used. The first-stage ranker can handle thousands of candidates and narrow them down so that the second-stage ranker only needs to process the highest-scoring 100 candidates.[7]

The first-stage ranker is trained to predict which items the second-stage ranker would place among its top results. This process, called knowledge distillation, allows the smaller first-stage model to imitate the decisions of the larger and more precise second-stage model while requiring much less computation. It also employs a two-tower neural network so that item embeddings can be calculated ahead of time and cached. Its design is very similar to the retrieval model, with the main differences being the size of the candidate pool and a different training objective.[8]

The second-stage ranker is applied to the remaining candidates to predict the probability of different types of engagement. This stage uses a multi-task, multi-label neural network, which is much heavier than the two-tower model. It lacks the benefit of caching because it takes in more powerful user-item interaction features that could not be used in the earlier two-tower stages, since including them would prevent the separation of item and user input data. However, this distinction is what makes the model more accurate in its predictions, while its computational cost is why the system waits until only the final 100 items remain before applying it. The model predicts probabilities for several possible actions, such as clicking, liking, or selecting “see less.” These probabilities are then combined using a value model, which applies different weights to create a final score used to order the remaining items. For example, increasing the weight assigned to likes would cause the ranking to place greater importance on the probability that a user will like a post.[9]

The Explore page also includes one final reranking stage in which the system can downrank potentially harmful content or apply rules that increase diversity, such as avoiding several posts from the same creator in a row.[10]

Illustration of algorithm optimization
A simplified view of Instagram’s ranking pipeline. User and post features are processed by a ranking model, which predicts the probability of different actions such as likes, comments, clicks, shares, and selecting “see less.” A value model then combines these predictions with different weights to produce a final ranking score. The probabilities and weights shown are illustrative, made-up values rather than actual Instagram parameters. Illustration created by the author with assistance from ChatGPT.

What Is the Algorithm Actually Optimizing?

Looking back at the value model used in the second-stage ranker, notice that it predicts specific Instagram actions such as likes, comments, and selecting “see less.” These are measurable behaviors, so they can be predicted by the model. User satisfaction or enjoyment with a given item, by contrast, is not directly observable. Because satisfaction is harder to observe, the system uses measurable actions as surrogate signals, or imperfect stand-ins for what the user may actually value. Different forms of engagement can then be assigned different weights. If liking a post is treated as a particularly strong signal of user satisfaction, for example, a larger weight can be applied to it.

This, however, assumes that the primary goal of social media platforms like Instagram is to maximize user happiness and satisfaction. That is not to say that such platforms don’t care about positive user experiences, but rather that the nature of social media as a multi-stakeholder system prevents user satisfaction from being the only goal. These platforms are businesses and must balance the interests of many different parties in order to stay in business. The main stakeholders are the users, the creators, the advertisers, and the company itself. The users want to receive enjoyable content, the creators want their posts to reach people and generate attention, and the advertisers want people to notice and respond to ads. The company itself wants revenue, continued use, growth, and manageable safety or reputation risks. These interests don’t depend solely on user satisfaction, but often converge on optimizing user engagement on the platform.[11] Higher engagement gives creators more exposure, gives advertisers more opportunities to reach users, and gives the platform more chances to generate revenue. It can also benefit users when engagement reflects genuine interest and helps the system find content they enjoy. However, this is highly platform-dependent, since each platform maintains its own ranking objectives. As a result, platforms use different algorithm designs, assign different weights to signals, and prioritize different qualities in their feeds. Depending on the platform, ranking systems may place different levels of importance on qualities such as timeliness, meaning how much their models prioritize newer content; novelty, meaning how much their models prioritize showing users unfamiliar content; specificity, meaning whether their models prioritize more general recommendations or niche recommendations matched to specific tastes; and usage intensity, meaning whether their models prioritize passive, duration-based viewing or active, engagement-based interactions such as comments, reactions, and shares.[12]

Illustration of engagement and enjoyment
A simplified view of the different stakeholders a social-media ranking system must balance. Users, creators, advertisers, and the platform itself each have different goals that can influence how content is ranked. Illustration created by the author with assistance from ChatGPT.

Engagement Is Not the Same as Enjoyment

Usage intensity, in particular, returns to the idea of user engagement being used as a proxy for user satisfaction. This, however, can lead the system to prioritize users’ revealed preferences over their stated preferences. Social media models interpret a user’s engagement with a content item as an indicator of interest and then use that revealed preference to update the user’s profile and recommend more or less similar content. The problem is that what users engage with does not always match what they would explicitly say they want to see, which the model usually has very little information about compared with the multitude of data it has on users’ revealed preferences.[13]

In a study first released in 2023 and formally published in 2025, Milli et al. conducted a preregistered independent audit of Twitter’s (now X’s) recommendation algorithm involving more than 800 active U.S. users. The study compared posts selected by Twitter’s engagement-based algorithm with the most recent posts from accounts each participant had chosen to follow. Participants then reported which posts they actually wanted to see and how the posts made them feel. This allowed the researchers to compare a feed selected using users’ revealed preferences with the posts users stated that they actually wanted to see.[14] The researchers found that the engagement-based feed amplified more emotionally charged, partisan, and hostile political content than the chronological feed. More importantly, users did not necessarily say that they preferred the political posts selected by the algorithm.[15] This suggests that supplying users with the content they will most enjoy is not necessarily the recommender system’s only or primary objective.

While this example was specific to X and mostly regarded political posts, it reveals a broader risk that could extend to other engagement-based social media platforms. Their recommender systems may have difficulty differentiating genuine enjoyment from negative reactions that still produce strong engagement, which the model then uses to influence future recommendations toward similar content, even when you do not actually like it.

The key takeaway here is that your feed is not controlled by you alone, and the system is not nearly as capable of understanding your feelings as its recommendations can make it seem. This does not mean that you should go and delete all social media (I certainly didn’t), but it is something you should be aware of as a member of not only an online but also a physical community. Enjoy your Reels, but don’t forget it’s a for-profit product and that the money ultimately comes from monetizing your attention.