Do Recommendation Algorithms Discover Our Preferences or Help Shape Them? - Exposure, Feedback, and Preference Formation

A person’s choices alter which cultural objects become more visible, and each new interaction feeds into the next selection, illustrating recommendation feedback without implying complete control.

 A person’s earlier choices shape what appears next, while each new response returns to the system as data, forming a continuous loop between preference, exposure, and behavior.



Do Recommendation Algorithms Discover Our Preferences or Help Shape Them? - Exposure, Feedback, and Preference Formation

Explore how recommendation, exposure, and user behavior form feedback loops that can blur the boundary between observed and underlying preferences.

Does an Algorithm Simply Observe What I Like?

Suppose I search for Nietzsche once. A system records an action. Depending on the service and its design, that action may become one of many signals used to rank, retrieve, or recommend later content. If Nietzsche-related material appears again and I select it, the system now possesses another observation. The sequence can continue.

At first, this seems like straightforward learning. I have a preference, my behavior reveals it, and the system becomes increasingly accurate at predicting what I will choose. But there is a scientific complication.

The system is not necessarily observing behavior in an environment independent of itself. By selecting, ranking, or recommending some items rather than others, it can help determine which options become available for the user to encounter. The next behavior may therefore occur partly inside an informational environment shaped by the previous prediction. This creates a feedback loop.

How Does a Recommendation Become New Data?

Collaborative filtering and related recommender methods use patterns in user interactions to estimate which items a person may prefer. Once a recommendation is presented, the user can accept it, ignore it, reject it, or interact with it in another measurable way. Some of those responses can subsequently become new data.

Sinha, Gleich, and Ramani studied precisely this problem. They describe how accepted recommendations can create feedback loops that iteratively influence later predictions of collaborative filtering systems. The important point is not merely that the system learns. It learns partly from behavior occurring after its own recommendations have entered the user's environment.

Consider a simplified sequence. A user has an existing interest. The system observes behavior related to that interest. It recommends an item. The user encounters the item because it was recommended. The user responds to it. That response becomes part of the data from which subsequent recommendations may be generated. The original preference has not disappeared. But the resulting dataset can now contain both traces of what the user previously preferred and traces of what occurred after the recommender intervened in exposure.

Can We Separate Preference from Recommendation?

This creates a measurement problem. If someone chooses an item after it has been recommended, how much of that choice reflects a preference that already existed, and how much reflects the fact that the item became visible?

Sinha and colleagues investigated whether the effects of recommender feedback could be separated from what they call intrinsic user preferences. Under specified assumptions, they developed a method for estimating the recommender system's influence on a user-item rating matrix and distinguishing recommended items from intrinsic preferences.

This does not mean that science has discovered a perfectly observable, permanent ‘true preference’ hidden inside every person. The term ‘intrinsic preference’ has a specific modeling role in their study. That limitation matters because human preferences can change for many reasons independent of recommendation systems. Experience, learning, social relationships, changing circumstances, deliberate exploration, and simple passage of time can all alter what a person chooses.

The scientific problem is therefore not to locate an eternally fixed preference. It is to determine whether observations produced after recommendation exposure can safely be treated as though they were independent of that exposure.

What Is Preference Pollution?

Adomavicius, Bockstedt, Curley, and Zhang describe a related problem as ‘preference pollution’. They model interaction between people and recommender systems as a continuous feedback loop containing pre-consumption and post-consumption phases. Before consumption, a system predicts preferences and provides recommendations. After consumption, the user may provide a rating or another form of feedback. That information can then contribute to subsequent recommendations. The difficulty arises when this post-consumption information is treated as ground-truth preference data even though the recommendation process itself may have influenced the decision that generated it.

The system can therefore face an unusual epistemic problem. It predicts behavior, helps structure the environment in which behavior occurs, observes the resulting behavior, and then uses that observation to improve later predictions.

This does not make the data false. It makes the origin of the data more complicated. The distinction is crucial. A click is an observation. What produced that click is a causal question.

Do Feedback Loops Make Everyone More Similar?

One possible consequence of repeated feedback is homogenization. Chaney, Stewart, and Engelhardt investigated this possibility by simulating recommender systems trained on data already affected by previous recommendations. In their simulations, algorithmic confounding increased homogenization of user behavior without increasing utility. This is an important result, but its evidential boundary must remain visible.

The study demonstrates a mechanism under simulated conditions. It does not establish that every real recommendation system makes human preferences converge or that every population exposed to personalization becomes culturally homogeneous. Real platforms contain many additional processes.

Users search deliberately. They follow links from outside a platform. Friends recommend material. People become bored. Interests change. Some users actively resist recommendations, while others rely heavily on them. The existence of a plausible feedback mechanism therefore does not determine the magnitude of its effect in every real environment.

What Happens on a Real Music Platform?

Research using Spotify provides a useful example of this distinction. Anderson, Maystre, Mehrotra, Anderson, and Lalmas analyzed music consumption on Spotify and examined the relationship between algorithmic recommendation and listening diversity. They found that algorithmically driven listening was associated with lower consumption diversity. Users whose listening became more diverse over time also tended to shift away from algorithmic consumption toward more organic consumption.

These findings might appear to show that recommendations narrow musical taste. But the researchers explicitly warn against that causal interpretation.

People who already possess narrower tastes may simply prefer algorithmically programmed listening. An observed association between algorithmic consumption and lower diversity therefore does not by itself reveal which direction the causal relationship runs. The study also included a randomized experiment, but its causal result concerned a different question: algorithmic recommendations were more effective for users whose existing listening diversity was lower.

This distinction between correlation and causation is essential. Observing what happens after recommendation is not automatically equivalent to knowing what the recommendation caused.

Does the Algorithm Discover or Shape Preference?

We can now return to the original question. Does a recommender system discover our preferences, or does it help shape them? The available evidence does not justify choosing only one side. Recommender systems can infer patterns from behavior that reflects preferences users already possess.

At the same time, recommendations affect exposure. Exposure creates opportunities for new behavior. Some of that behavior can become new training data. The system is therefore neither a perfectly passive observer nor necessarily a sovereign producer of preference. It participates in a coupled process.

User preferences affect behavior. Behavior affects recommendations. Recommendations affect exposure. Exposure creates new opportunities for behavior. New behavior can affect later recommendations. The scientific difficulty lies precisely in separating these relations. When the system's outputs help generate its future inputs, prediction and observation are no longer completely independent stages.

What Does the Algorithm Actually Know About Me?

A recommender system does not directly observe a complete human preference. It observes data. Searches, ratings, clicks, listening histories, purchases, skips, viewing times, and other interactions can function as signals, depending on the particular system. From these signals, models estimate patterns useful for prediction or ranking.

But an estimated preference is not identical to the person. A person may search for something out of curiosity and never return to it. Another may repeatedly consume something while wishing to stop. Someone else may discover a subject accidentally and continue studying it for decades. The same visible behavior can emerge from different histories and intentions.

For this reason, the most precise scientific question is not simply whether algorithms ‘know what we like’. It is what can legitimately be inferred about preference from behavioral data generated inside environments that recommendation systems themselves partly organize. That question remains open to measurement, experimentation, and better causal identification.

This also changes the way we should understand the familiar recommendation loop. The algorithm does not merely look backward at a completed record of who we were. Its predictions can also become part of the environment in which the next record is made.



References

Sinha, Ayan, David F. Gleich, and Karthik Ramani. “Deconvolving Feedback Loops in Recommender Systems.” ‘Advances in Neural Information Processing Systems’, vol. 29, 2016, pp. 3243–3251.

This study was used to establish the feedback loop between recommendation and subsequent user behavior and to examine the distinction between recommender-system influence and intrinsic user preference.

Adomavicius, Gediminas, Jesse C. Bockstedt, Shawn P. Curley, and Jingjing Zhang. “Recommender Systems, Ground Truth, and Preference Pollution.” ‘AI Magazine’, vol. 43, no. 2, 2022, pp. 177–189. DOI: 10.1002/aaai.12055.

This article was used to examine the continuous feedback loop between pre-consumption recommendation and post-consumption feedback and the resulting problem of preference pollution in ground-truth data.

Chaney, Allison J. B., Brandon M. Stewart, and Barbara E. Engelhardt. “How Algorithmic Confounding in Recommendation Systems Increases Homogeneity and Decreases Utility.” ‘Proceedings of the 12th ACM Conference on Recommender Systems’, 2018, pp. 224–232. DOI: 10.1145/3240323.3240370.

This study was used to examine, specifically under simulation, how training on data already affected by recommendations can create algorithmic confounding and increase behavioral homogenization.

Anderson, Ashton, Lucas Maystre, Rishabh Mehrotra, Ian Anderson, and Mounia Lalmas. “Algorithmic Effects on the Diversity of Consumption on Spotify.” ‘Proceedings of The Web Conference 2020’, 2020, pp. 2155–2165. DOI: 10.1145/3366423.3380281.

This study was used to examine real-platform evidence concerning algorithmic consumption and listening diversity while preserving the authors’ explicit distinction between observed associations and causal conclusions.


한 사람의 이전 선택이 다음에 보이는 것을 형성하고, 새로운 반응은 다시 데이터가 되어 시스템으로 돌아가면서 선호, 노출과 행동 사이에 지속적인 순환을 만듭니다.


추천 알고리즘은 우리의 취향을 발견하는가, 형성에도 관여하는가? - 노출, 피드백, 선호 형성

추천, 노출과 사용자 행동이 어떻게 피드백 루프를 형성하여 관찰된 선호와 기저 선호의 경계를 흐릴 수 있는지 살펴봅니다.


알고리즘은 단순히 내가 좋아하는 것을 관찰할까?

제가 니체를 한 번 검색했다고 가정해 보겠습니다. 시스템은 하나의 행동을 기록합니다. 서비스와 설계 방식에 따라 그 행동은 이후 콘텐츠의 순위를 정하거나 검색하거나 추천하는 데 사용되는 여러 신호 가운데 하나가 될 수 있습니다. 니체와 관련된 자료가 다시 나타나고 제가 그것을 선택하면 시스템에는 또 하나의 관찰값이 생깁니다. 이 과정은 계속될 수 있습니다.

처음에는 단순한 학습처럼 보입니다. 저에게 취향이 있고, 행동이 그것을 드러내며, 시스템은 제가 무엇을 선택할지 점점 더 정확하게 예측합니다. 그러나 여기에는 과학적인 문제가 하나 있습니다.

시스템은 반드시 자신과 독립된 환경에서 일어나는 행동만을 관찰하는 것이 아닙니다. 일부 항목을 다른 항목보다 먼저 선택하거나 높은 순위에 놓거나 추천함으로써 사용자가 어떤 선택지를 접할 수 있는지 결정하는 데 관여할 수 있습니다. 따라서 다음 행동은 이전의 예측에 의해 부분적으로 구성된 정보 환경 안에서 발생할 수 있습니다. 여기에서 피드백 루프가 생깁니다.


추천은 어떻게 새로운 데이터가 될까?

협업 필터링과 관련 추천 기법은 사용자의 상호작용에서 나타나는 패턴을 이용하여 한 사람이 어떤 항목을 선호할 가능성이 있는지를 추정합니다. 추천이 제시되면 사용자는 그것을 받아들이거나, 무시하거나, 거부하거나, 다른 측정 가능한 방식으로 상호작용할 수 있습니다. 그러한 반응 가운데 일부는 이후 새로운 데이터가 될 수 있습니다.

신하, 글라이히와 라마니는 바로 이 문제를 연구했습니다. 이들은 사용자가 받아들인 추천이 피드백 루프를 만들고, 그 루프가 협업 필터링 시스템의 이후 예측에 반복적으로 영향을 미칠 수 있는 구조를 설명합니다. 중요한 것은 단순히 시스템이 학습한다는 사실이 아닙니다. 시스템은 자신의 추천이 사용자의 환경에 들어간 뒤 발생한 행동에서도 부분적으로 학습합니다.

단순화한 과정을 생각해 보겠습니다. 사용자에게 기존의 관심이 있습니다. 시스템은 그 관심과 관련된 행동을 관찰합니다. 시스템이 하나의 항목을 추천합니다. 사용자는 추천되었기 때문에 그 항목을 접합니다. 사용자가 반응합니다. 그 반응은 이후의 추천을 생성하는 데이터의 일부가 될 수 있습니다. 처음의 선호가 사라진 것은 아닙니다. 그러나 그 결과 만들어진 데이터에는 사용자가 이전부터 선호했던 것의 흔적과 추천 시스템이 노출에 개입한 뒤 일어난 일의 흔적이 함께 들어갈 수 있습니다.


선호와 추천을 분리할 수 있을까?

여기에서 측정의 문제가 생깁니다. 어떤 사람이 추천받은 항목을 선택했다면, 그 선택 가운데 얼마만큼이 이미 존재하던 선호를 반영하고 얼마만큼이 그 항목이 눈앞에 나타났기 때문에 발생한 것일까요?

신하와 동료들은 추천 시스템의 피드백 효과를 이들이 ‘내재적 사용자 선호’라고 부르는 것과 분리할 수 있는지를 연구했습니다. 이들은 일정한 가정 아래 추천 시스템이 사용자–항목 평점 행렬에 미친 영향을 추정하고, 추천된 항목과 내재적 선호를 구분하는 방법을 개발했습니다.

그렇다고 과학이 모든 인간의 내부에 숨어 있는 완벽하게 관찰 가능한 영구적인 ‘진짜 취향’을 발견했다는 뜻은 아닙니다. ‘내재적 선호’라는 용어는 해당 연구의 모형 안에서 특정한 역할을 합니다. 이러한 한계는 중요합니다. 인간의 선호는 추천 시스템과 관계없는 수많은 이유로도 변할 수 있기 때문입니다. 경험, 학습, 사회적 관계, 환경의 변화, 의도적인 탐색과 단순한 시간의 흐름도 사람이 선택하는 것을 변화시킬 수 있습니다.

따라서 과학적 문제는 영원히 고정된 하나의 선호를 찾아내는 것이 아닙니다. 추천에 노출된 뒤 만들어진 관찰값을 그 노출과 독립적으로 만들어진 것처럼 취급해도 되는지를 판단하는 것이 문제입니다.


‘선호 오염’이란 무엇일까?

아도마비시우스, 복스테트, 컬리와 장은 이와 관련된 문제를 ‘선호 오염’이라고 설명합니다. 이들은 인간과 추천 시스템의 상호작용을 소비 전 단계와 소비 후 단계가 포함된 지속적인 피드백 루프로 봅니다. 소비 전에는 시스템이 선호를 예측하여 추천합니다. 소비 후에는 사용자가 평점이나 다른 형태의 피드백을 제공할 수 있습니다. 그 정보는 다시 이후의 추천에 사용될 수 있습니다. 문제는 소비 후 정보가 추천 과정 자체의 영향을 받은 결정에서 만들어졌을 가능성이 있는데도 그것을 실제 선호를 나타내는 ‘ground truth’ 데이터로 취급할 때 발생합니다.

따라서 시스템은 특이한 인식론적 문제에 직면할 수 있습니다. 행동을 예측하고, 그 행동이 일어나는 환경의 구성에 관여하고, 그 결과로 나타난 행동을 관찰한 뒤, 그 관찰값으로 이후의 예측을 개선합니다.

그렇다고 데이터가 거짓이라는 뜻은 아닙니다. 데이터가 만들어진 기원이 더 복잡하다는 뜻입니다. 이 구별은 중요합니다. 클릭은 관찰값입니다. 그 클릭을 무엇이 만들어 냈는가는 인과관계의 문제입니다.


피드백 루프는 모두를 비슷하게 만들까?

반복적인 피드백에서 나타날 수 있는 결과 가운데 하나는 동질화입니다. 체이니, 스튜어트와 엥겔하르트는 이전 추천의 영향을 이미 받은 데이터로 학습하는 추천 시스템을 시뮬레이션하여 이러한 가능성을 연구했습니다. 그들의 시뮬레이션에서는 알고리즘적 교란이 효용을 증가시키지 않으면서 사용자 행동의 동질화를 증가시켰습니다. 이 결과는 중요하지만 근거의 경계를 분명하게 유지해야 합니다.

이 연구는 시뮬레이션 조건에서 하나의 메커니즘을 보여 줍니다. 모든 실제 추천 시스템이 인간의 취향을 수렴시키거나 개인화에 노출된 모든 집단이 문화적으로 동질화된다는 사실을 확립하지는 않습니다. 실제 플랫폼에는 훨씬 많은 과정이 존재합니다.

사용자는 의도적으로 검색합니다. 플랫폼 밖의 링크를 따라갑니다. 친구가 자료를 추천합니다. 사람들은 싫증을 느끼고 관심사는 변합니다. 추천에 적극적으로 저항하는 사용자도 있고 추천에 크게 의존하는 사용자도 있습니다. 따라서 가능한 피드백 메커니즘이 존재한다는 사실만으로 모든 실제 환경에서 그 효과의 크기를 결정할 수는 없습니다.


실제 음악 플랫폼에서는 어떤 일이 일어날까?

Spotify를 이용한 연구는 이 구별을 보여 주는 유용한 사례입니다. 앤더슨, 메이스트르, 메로트라, 앤더슨과 랄마스는 Spotify의 음악 소비를 분석하고 알고리즘 추천과 청취 다양성의 관계를 조사했습니다. 이들은 알고리즘에 의해 이루어진 청취가 낮은 소비 다양성과 연관되어 있음을 발견했습니다. 시간이 지나면서 청취가 더 다양해진 사용자들은 알고리즘에 의한 소비에서 벗어나 자발적인 소비를 늘리는 경향도 보였습니다.

이러한 결과는 추천이 음악적 취향을 좁힌다는 의미로 보일 수 있습니다. 그러나 연구자들은 그러한 인과적 해석을 명시적으로 경계합니다.

원래 더 좁은 취향을 가진 사람들이 알고리즘으로 구성된 청취 방식을 더 선호할 수도 있습니다. 따라서 알고리즘 소비와 낮은 다양성 사이의 관찰된 연관성만으로 인과관계가 어느 방향으로 작동하는지 알 수 없습니다. 이 연구에는 무작위 실험도 포함되어 있지만 인과적으로 확인한 질문은 다릅니다. 알고리즘 추천은 기존의 청취 다양성이 낮은 사용자에게 더 효과적이었습니다.

상관관계와 인과관계를 구별하는 것은 필수적입니다. 추천 이후에 어떤 일이 일어났다는 사실을 관찰하는 것과 추천이 무엇을 일으켰는지를 아는 것은 같은 일이 아닙니다.


알고리즘은 취향을 발견할까, 형성할까?

이제 처음의 질문으로 돌아갈 수 있습니다. 추천 시스템은 우리의 취향을 발견하는 것일까요, 아니면 그것을 형성하는 데 관여하는 것일까요? 현재의 근거만으로 어느 한쪽만 선택하는 것은 정당하지 않습니다. 추천 시스템은 사용자가 이미 지닌 선호를 반영하는 행동에서 패턴을 추론할 수 있습니다.

동시에 추천은 노출에 영향을 줍니다. 노출은 새로운 행동이 발생할 기회를 만듭니다. 그 행동 가운데 일부는 새로운 학습 데이터가 될 수 있습니다. 따라서 시스템은 완전히 수동적인 관찰자도 아니며 반드시 선호를 지배적으로 생산하는 존재도 아닙니다. 인간과 시스템은 서로 연결된 과정에 참여합니다.

사용자의 선호는 행동에 영향을 줍니다. 행동은 추천에 영향을 줍니다. 추천은 노출에 영향을 줍니다. 노출은 새로운 행동의 기회를 만듭니다. 새로운 행동은 다시 이후의 추천에 영향을 줄 수 있습니다. 과학적 어려움은 바로 이러한 관계들을 분리하는 데 있습니다. 시스템의 출력이 미래의 입력을 생성하는 데 관여한다면 예측과 관찰은 더 이상 완전히 독립된 단계가 아닙니다.


알고리즘은 실제로 나에 대해 무엇을 알까?

추천 시스템이 인간의 완전한 취향을 직접 관찰하는 것은 아닙니다. 시스템이 관찰하는 것은 데이터입니다. 특정 시스템의 설계에 따라 검색, 평점, 클릭, 청취 기록, 구매, 건너뛰기, 시청 시간과 다른 상호작용이 신호로 사용될 수 있습니다. 모형은 이러한 신호에서 예측이나 순위 결정에 유용한 패턴을 추정합니다.

그러나 추정된 선호와 인간 그 자체는 동일하지 않습니다. 호기심 때문에 무언가를 한 번 검색하고 다시는 찾지 않는 사람도 있습니다. 어떤 것을 반복해서 소비하면서도 그만두고 싶어하는 사람도 있습니다. 우연히 하나의 주제를 발견한 뒤 수십 년 동안 연구하는 사람도 있습니다. 겉으로 동일한 행동도 서로 다른 역사와 의도에서 나올 수 있습니다.

따라서 과학적으로 더 정확한 질문은 알고리즘이 단순히 ‘우리가 무엇을 좋아하는지 아는가’가 아닙니다. '추천 시스템 자체가 부분적으로 구성한 환경 안에서 만들어진 행동 데이터로부터 선호에 관해 무엇을 정당하게 추론할 수 있는가'가 문제입니다. 이 질문은 측정, 실험과 더 나은 인과관계 식별을 통해 계속 검토되어야 합니다.

이 질문은 또한 익숙한 추천의 순환을 바라보는 방식을 바꿉니다. 알고리즘은 이미 완성된 과거의 나에 관한 기록을 단순히 뒤돌아보는 것만은 아닙니다. 알고리즘의 예측은 다음 기록이 만들어지는 환경의 일부가 될 수도 있습니다.


참고문헌

아얀 신하, 데이비드 F. 글라이히, 카르티크 라마니. “추천 시스템의 피드백 루프 분리.” ‘Advances in Neural Information Processing Systems’, 제29권, 2016, 3243–3251쪽.

추천과 이후의 사용자 행동 사이에서 발생하는 피드백 루프를 확인하고, 추천 시스템의 영향과 연구에서 말하는 사용자의 내재적 선호를 구별하는 문제를 검토하는 데 사용했습니다. 논문 제목은 이해를 돕기 위해 한글로 옮겼습니다.


게디미나스 아도마비치우스, 제시 C. 복스테트, 숀 P. 컬리, 징징 장. “추천 시스템, 정답 데이터와 선호 오염.” ‘AI Magazine’, 제43권 제2호, 2022, 177–189쪽. DOI: 10.1002/aaai.12055.

소비 이전의 추천과 소비 이후의 피드백이 이어지는 지속적인 피드백 루프를 검토하고, 그 과정에서 정답으로 취급되는 선호 데이터에 ‘선호 오염’ 문제가 발생할 수 있다는 논의를 확인하는 데 사용했습니다. 논문 제목은 이해를 돕기 위해 한글로 옮겼습니다.


앨리슨 J. B. 체이니, 브랜던 M. 스튜어트, 바버라 E. 엥겔하르트. “추천 시스템의 알고리즘 교란은 어떻게 동질성을 높이고 효용을 낮추는가.” ‘제12회 ACM 추천 시스템 학술대회 논문집’, 2018, 224–232쪽. DOI: 10.1145/3240323.3240370.

추천의 영향을 이미 받은 데이터를 다시 학습할 때 알고리즘 교란과 행동 동질화가 발생할 수 있다는 결과를 검토하는 데 사용했습니다. 이 결과가 시뮬레이션에서 도출되었다는 근거 범위를 유지하기 위해 사용했습니다. 논문 제목은 이해를 돕기 위해 한글로 옮겼습니다.


애슈턴 앤더슨, 루카스 메이스트르, 리샤브 메흐로트라, 이언 앤더슨, 무니아 랄마스. “스포티파이 소비 다양성에 대한 알고리즘의 효과.” ‘The Web Conference 2020 논문집’, 2020, 2155–2165쪽. DOI: 10.1145/3366423.3380281.

실제 음악 플랫폼에서 알고리즘 기반 소비와 청취 다양성의 관계를 검토하면서, 관찰된 상관관계와 인과적 결론을 구분하는 근거로 사용했습니다. 또한 무작위 실험에서 기존 청취 다양성이 낮은 사용자에게 알고리즘 추천이 더 효과적이었다는 별도의 인과 결과를 구분하는 데 사용했습니다. 논문 제목은 이해를 돕기 위해 한글로 옮겼습니다.



Previous Episode - How Does the Brain Distinguish Reality from Imagination? - Perception, Dreams, and Reality Monitoring

Related Label - Science

Next Episode - 

Comments

Popular posts from this blog

Human Story Lab Introduction (Human Story Lab 소개)

The Odyssey Introduction – What Does a Human Being Seek After War?

The Odyssey Episode 3 - Circe and the Island of Magic