Does an Algorithm Tell a Story About Who I Am? - Digital Traces, Prediction, and the Algorithmic Self
A person moves forward while digital traces from earlier choices remain behind, forming an incomplete portrait that follows but never fully becomes the person.
Does an Algorithm Tell a Story About Who I Am? - Digital Traces, Prediction, and the Algorithmic Self
Explore how digital traces become predictions about who we are, revealing the gap between algorithmic profiles and changing human lives.
What Does a Single Search Say About Me?
Suppose I search for Nietzsche today. I may be studying philosophy, checking a quotation, writing an article, or simply following a moment of curiosity. Tomorrow, my attention may move to astronomy, an ancient language, a novel, or something entirely unrelated.
For me, the search belongs to a particular moment and has a particular reason. A digital system, however, does not need access to that private reason in order to register an observable action. What remains available to computational processing is a trace: a query, a click, a viewing history, a rating, a purchase, or another recorded interaction.
This distinction matters because a trace is evidence of an action, not a complete explanation of the person who performed it. One search for Nietzsche establishes that a search occurred. By itself, it does not establish why the person searched, how important Nietzsche is to that person, or whether the interest will continue tomorrow.
Yet digital systems rarely operate from one isolated trace. Multiple traces can be collected, classified, and used to make inferences about interests, preferences, and likely future behavior. Under the European Union’s General Data Protection Regulation, profiling explicitly includes automated processing used to analyze or predict aspects such as a person’s preferences, interests, and behavior.
A momentary action can therefore become part of a longer computational representation of a person.
How Does a Trace Become a Version of Me?
John Cheney-Lippold described one form of this process as ‘algorithmic identity’. His analysis focused on how web analytics and marketing systems can infer categories of identity from patterns in users’ internet activity.
The important distinction is between a person and a classification produced about that person. An inferred category is not the person itself. It is a computationally useful representation constructed from available data according to particular methods and purposes.
This is not unique to algorithms. Human beings also understand one another through incomplete evidence. We remember actions, interpret words, notice habits, and construct expectations. A biographer likewise selects a limited number of events from a life that contained far more moments than any book could preserve.
But the comparison has limits. A computational profile does not become a narrator with human consciousness merely because it organizes traces into categories or predictions. It does not remember a life as a person remembers a life, nor does it necessarily understand why an action mattered.
The narrative comparison becomes useful at a structural level. A story cannot contain every event. It selects. It connects. It gives some traces greater relevance than others. A digital profile also depends on selection: which behaviors are recorded, which variables are retained, which categories are defined, and which patterns are treated as useful for prediction.
The result is not ‘me’. It is one data-based version of me produced for a particular computational purpose.
What Happens When the Past Predicts the Future?
The temporal problem is more difficult. Human beings change, while records of earlier behavior can remain available for later processing.
Yesterday’s curiosity can therefore participate in today’s prediction. If I respond to what is subsequently recommended, that response may itself become another recorded interaction.
Research on recommender systems has examined precisely this kind of feedback loop. Ayan Sinha, David Gleich, and Karthik Ramani studied how accepting recommendations can influence the data from which later recommendations are generated. Their work also addressed whether the influence of recommendation could be distinguished from what they called intrinsic user preference.
This distinction exposes a fundamental difficulty. Once a system has influenced what a person encounters, subsequent behavior is no longer simply an untouched measurement of what the person would have chosen without that exposure.
Allison Chaney, Brandon Stewart, and Barbara Engelhardt described a related problem as algorithmic confounding. In simulations, they found that repeatedly training recommendation systems on behavior already affected by previous recommendations could increase homogenization in user behavior without increasing utility.
These findings do not establish that every recommendation system traps every person in a fixed identity. Systems differ, users respond differently, and simulated effects should not automatically be treated as universal observations of human behavior.
They establish something narrower but important: the past can be used to predict the future, while responses to those predictions can become part of the next round of data.
Can a Profile Become an Outdated Story?
Now the narrative problem becomes clearer.
A human life contains discontinuity. People abandon interests, encounter unfamiliar ideas, change professions, revise beliefs, acquire new tastes, and sometimes become interested in things that their previous behavior could not have predicted.
A representation built from past traces necessarily begins behind the present. This does not make prediction useless. Past behavior can contain valuable information. But predictive usefulness and complete representation are different claims.
A biography written at age twenty cannot contain the life lived at forty. In a similar way, a profile derived from previous behavior cannot contain choices that have not yet occurred. It can estimate them only from the information and model available to it.
This is where accuracy and discovery can come into tension. Recommender-system research has long recognized that predictive accuracy is not the only relevant measure of recommendation quality. Coverage and serendipity have also been proposed as important dimensions because useful discovery can involve encountering something that is not merely an obvious repetition of previous choices.
A system that perfectly repeats the recognizable pattern of my past might therefore describe one aspect of me very well while leaving another aspect unexplored: my capacity to become interested in something new.
Who Am I Beyond My Digital Traces?
The question is not whether digital traces are false. Many of them record things we actually did. I searched, clicked, watched, bought, paused, returned, or ignored.
The question is what kind of truth those actions can support.
A record can establish an event without fully establishing its meaning. A sequence of searches can reveal recurring behavior without explaining every motive behind it. A prediction can be statistically useful without becoming a complete account of the person whose behavior it predicts.
This is why ‘algorithmic self’ is best treated carefully. The profile is not another conscious self living inside a machine. It is a representation generated from data, classifications, and predictive procedures.
Yet that representation matters because it can participate in determining what becomes visible next. The version of me inferred from yesterday may help organize the environment presented to me today.
Narrative therefore gives us a useful final question. Every story about a person is selective, but human beings are not finished stories. We continue to act beyond what has already been recorded.
Perhaps the most important difference between a person and an algorithmic profile lies there. The profile must begin with traces that already exist. A human life can still produce the trace that did not exist yesterday.
The question is therefore not simply, “Does the algorithm know me?”
A more precise question is, “How much of the person I am becoming can ever be contained in a story constructed from the person I have already been?”
References
Cheney-Lippold, John. “A New Algorithmic Identity: Soft Biopolitics and the Modulation of Control.” ‘Theory, Culture & Society’, vol. 28, no. 6, 2011, pp. 164–181. DOI: 10.1177/0263276411424420.
This article was used to establish the concept of algorithmic identity and to examine how categories about users can be inferred from patterns in online activity.
European Union. ‘Regulation (EU) 2016/679 of the European Parliament and of the Council’, 2016, Article 4(4).
This official legal text was used to verify the definition of profiling as automated processing used to evaluate, analyze, or predict personal aspects including preferences, interests, and behavior.
Sinha, Ayan, David F. Gleich, and Karthik Ramani. “Deconvolving Feedback Loops in Recommender Systems.” ‘Advances in Neural Information Processing Systems’, vol. 29, 2016, pp. 3243–3251.
This study was used to examine feedback between recommendations and subsequent user behavior and the problem of distinguishing recommendation influence from intrinsic preference.
Chaney, Allison J. B., Brandon M. Stewart, and Barbara E. Engelhardt. “How Algorithmic Confounding in Recommendation Systems Increases Homogeneity and Decreases Utility.” ‘Proceedings of the 12th ACM Conference on Recommender Systems’, 2018, pp. 224–232. DOI: 10.1145/3240323.3240370.
This study was used to examine, through simulation, how repeatedly learning from behavior already influenced by recommendations can produce algorithmic confounding and greater behavioral homogeneity.
Ge, Mouzhi, Carla Delgado-Battenfeld, and Dietmar Jannach. “Beyond Accuracy: Evaluating Recommender Systems by Coverage and Serendipity.” ‘Proceedings of the Fourth ACM Conference on Recommender Systems’, 2010, pp. 257–260. DOI: 10.1145/1864708.1864761.
This study was used to support the distinction between predictive accuracy and broader recommendation qualities such as coverage and serendipity.
한 사람이 앞으로 나아가는 동안 이전 선택의 디지털 흔적은 뒤에 남아, 그 사람을 따라가지만 결코 그 사람 전체가 되지는 못하는 불완전한 초상을 형성합니다.
알고리즘은 나에 관한 이야기를 만드는가? - 디지털 흔적, 예측, 알고리즘적 자아
디지털 흔적이 우리가 누구인지에 관한 예측으로 바뀌는 과정을 통해 알고리즘적 프로필과 변화하는 인간의 삶 사이의 간극을 살펴봅니다.
한 번의 검색은 나에 대해 무엇을 말할 수 있을까?
오늘 제가 니체를 검색했다고 가정해 보겠습니다. 철학을 공부하고 있을 수도 있고, 인용문을 확인하고 있을 수도 있으며, 글을 쓰고 있거나 단순히 순간적인 호기심을 따라갔을 수도 있습니다. 내일이면 관심이 천문학이나 고대 언어, 소설 또는 전혀 관계없는 무엇인가로 옮겨갈 수도 있습니다.
저에게 그 검색은 특정한 순간에 특정한 이유로 이루어진 행동입니다. 그러나 디지털 시스템은 관찰할 수 있는 행동을 기록하기 위해 그 사적인 이유에 접근할 필요가 없습니다. 컴퓨터 처리의 대상으로 남는 것은 검색어, 클릭, 시청 기록, 평가, 구매 또는 그 밖의 기록된 상호작용이라는 흔적입니다.
이 구분이 중요한 이유는 흔적이 행동의 증거이지, 그 행동을 한 사람에 관한 완전한 설명은 아니기 때문입니다. 니체를 한 번 검색했다는 것은 검색이라는 행동이 일어났음을 보여 줍니다. 그것만으로는 왜 검색했는지, 그 사람에게 니체가 얼마나 중요한지, 내일도 그 관심이 계속될지를 확정할 수 없습니다.
그러나 디지털 시스템은 대개 하나의 고립된 흔적만으로 작동하지 않습니다. 여러 흔적을 수집하고 분류하여 관심, 선호와 앞으로 나타날 가능성이 있는 행동을 추론하는 데 사용할 수 있습니다. 유럽연합의 일반개인정보보호법은 프로파일링에 개인의 선호, 관심과 행동 등을 분석하거나 예측하기 위한 자동화된 처리를 명시적으로 포함합니다.
따라서 순간적인 행동은 한 사람에 관한 더 장기적인 컴퓨터상의 표상을 구성하는 일부가 될 수 있습니다.
흔적은 어떻게 나의 한 가지 버전이 될까?
존 체니-리폴드는 이러한 과정의 한 형태를 ‘알고리즘적 정체성’이라고 설명했습니다. 그의 분석은 웹 분석과 마케팅 시스템이 사용자의 인터넷 활동 패턴으로부터 정체성의 범주를 어떻게 추론할 수 있는지에 초점을 맞추었습니다.
여기서 중요한 것은 인간과 그 인간에 관해 만들어진 분류를 구별하는 것입니다. 추론된 범주는 인간 그 자체가 아닙니다. 이용할 수 있는 데이터로부터 특정한 방법과 목적에 따라 만들어진, 컴퓨터 처리에 유용한 표상입니다.
이러한 방식이 알고리즘에만 존재하는 것은 아닙니다. 인간도 불완전한 증거를 통해 서로를 이해합니다. 우리는 행동을 기억하고, 말을 해석하고, 습관을 발견하며, 앞으로의 행동을 예상합니다. 전기 작가 역시 어떤 삶에 실제로 존재했던 수많은 순간 가운데 한 권의 책이 담을 수 있는 제한된 사건만을 선택합니다.
그러나 이 비교에는 한계가 있습니다. 컴퓨터상의 프로필이 흔적을 범주나 예측으로 조직한다고 해서 인간의 의식을 가진 서술자가 되는 것은 아닙니다. 인간이 자신의 삶을 기억하는 방식으로 삶을 기억하지 않으며, 어떤 행동이 그 사람에게 왜 중요했는지를 반드시 이해하는 것도 아닙니다.
서사와의 비교가 유용해지는 것은 구조적인 차원입니다. 이야기는 모든 사건을 담을 수 없습니다. 선택하고, 연결하고, 어떤 흔적에는 다른 흔적보다 더 큰 중요성을 부여합니다. 디지털 프로필 역시 무엇을 선택하느냐에 의존합니다. 어떤 행동을 기록하는지, 어떤 변수를 보존하는지, 어떤 범주를 정의하는지, 어떤 패턴을 예측에 유용하다고 판단하는지가 중요합니다.
그 결과는 ‘나’가 아닙니다. 특정한 컴퓨터 처리의 목적을 위해 만들어진 데이터 기반의 ‘나의 한 가지 버전’입니다.
과거가 미래를 예측하면 무슨 일이 일어날까?
시간의 문제는 더욱 어렵습니다. 인간은 변하지만 과거 행동의 기록은 이후의 처리에 이용할 수 있는 상태로 남을 수 있습니다.
따라서 어제의 호기심이 오늘의 예측에 관여할 수 있습니다. 이후 추천된 것에 제가 반응한다면 그 반응 자체도 또 하나의 기록된 상호작용이 될 수 있습니다.
추천 시스템 연구에서는 바로 이러한 종류의 피드백 루프를 다루어 왔습니다. 아얀 신하, 데이비드 글라이히와 카르티크 라마니는 추천을 받아들이는 행동이 이후의 추천을 생성하는 데이터에 어떻게 영향을 줄 수 있는지 연구했습니다. 이들의 연구는 추천의 영향과 이른바 사용자의 ‘내재적 선호’를 구별할 수 있는지도 다루었습니다.
이러한 구분은 근본적인 어려움을 드러냅니다. 시스템이 어떤 사람이 무엇을 접하는지에 이미 영향을 준 뒤라면, 그 이후의 행동을 그러한 노출이 없었을 때 그 사람이 선택했을 것을 그대로 측정한 결과라고 단순하게 볼 수 없습니다.
앨리슨 체이니, 브랜던 스튜어트와 바버라 엥겔하르트는 이와 관련된 문제를 ‘알고리즘 교란’으로 설명했습니다. 이들은 시뮬레이션에서 이전의 추천에 이미 영향을 받은 행동을 다시 사용해 추천 시스템을 반복적으로 학습시키면 효용의 증가 없이 사용자 행동의 동질화가 증가할 수 있음을 확인했습니다.
이러한 연구 결과가 모든 추천 시스템이 모든 인간을 고정된 정체성 안에 가둔다는 것을 입증하지는 않습니다. 시스템마다 차이가 있고 사용자의 반응도 다르며, 시뮬레이션에서 나타난 효과를 인간 행동에 관한 보편적인 관찰 결과로 자동 확대해서도 안 됩니다.
이 연구들이 보여 주는 범위는 더 좁지만 중요합니다. 과거가 미래를 예측하는 데 이용될 수 있으며, 그 예측에 대한 반응이 다시 다음 데이터의 일부가 될 수 있다는 것입니다.
프로필은 낡은 이야기가 될 수 있을까?
이제 서사의 문제가 더욱 분명해집니다.
인간의 삶에는 불연속성이 존재합니다. 사람은 관심을 버리고, 낯선 생각을 만나고, 직업을 바꾸고, 믿음을 수정하고, 새로운 취향을 가지며, 때로는 과거의 행동으로 예측하기 어려웠던 것에 관심을 갖게 됩니다.
과거의 흔적으로 만들어진 표상은 필연적으로 현재보다 앞선 시점에서 출발합니다. 그렇다고 예측이 쓸모없다는 뜻은 아닙니다. 과거의 행동에는 가치 있는 정보가 포함될 수 있습니다. 그러나 예측에 유용하다는 것과 인간을 완전하게 표상한다는 것은 서로 다른 주장입니다.
스무 살에 쓰인 전기는 마흔 살에 살아갈 삶을 담을 수 없습니다. 마찬가지로 이전 행동으로부터 만들어진 프로필에는 아직 이루어지지 않은 선택이 들어 있을 수 없습니다. 이용할 수 있는 정보와 모형을 통해 그것을 추정할 수 있을 뿐입니다.
바로 이 지점에서 정확성과 발견 사이에 긴장이 생길 수 있습니다. 추천 시스템 연구에서는 오래전부터 예측 정확도만이 추천의 품질을 평가하는 유일한 기준이 아니라는 점을 다루어 왔습니다. 유용한 발견에는 이전 선택의 뻔한 반복이 아닌 것을 만나는 경험도 포함될 수 있기 때문에 범위와 세렌디피티 역시 중요한 평가 차원으로 제안되었습니다.
따라서 과거의 인식 가능한 패턴을 완벽하게 반복하는 시스템은 나의 한 측면을 매우 정확하게 묘사하면서도 또 다른 측면을 탐색하지 못할 수 있습니다. 바로 새로운 것에 관심을 가질 수 있는 능력입니다.
디지털 흔적 너머의 나는 누구일까?
문제는 디지털 흔적이 거짓이냐는 것이 아닙니다. 그 가운데 많은 것은 우리가 실제로 한 행동을 기록합니다. 저는 검색하고, 클릭하고, 보고, 구매하고, 멈추고, 다시 돌아가거나 무시했습니다.
문제는 그러한 행동이 어떤 종류의 진실을 뒷받침할 수 있는가입니다.
기록은 사건이 일어났음을 확인하면서도 그 의미 전체를 확정하지 못할 수 있습니다. 일련의 검색 기록은 반복되는 행동을 보여 주면서도 그 뒤에 있는 모든 동기를 설명하지는 못합니다. 예측은 통계적으로 유용하면서도 그 대상이 되는 인간에 관한 완전한 설명이 아닐 수 있습니다.
따라서 ‘알고리즘적 자아’라는 표현은 신중하게 사용해야 합니다. 프로필은 기계 안에서 살아가는 또 하나의 의식적인 자아가 아닙니다. 데이터, 분류와 예측 절차로부터 만들어진 표상입니다.
그러나 그 표상은 다음에 무엇이 보이게 될지를 결정하는 데 관여할 수 있기 때문에 중요합니다. 어제의 나로부터 추론된 나의 버전이 오늘 제시되는 환경을 조직하는 데 이용될 수 있습니다.
그러므로 서사는 마지막에 유용한 질문을 하나 던집니다. 한 사람에 관한 모든 이야기는 선택적이지만 인간은 완성된 이야기가 아닙니다. 우리는 이미 기록된 것 너머에서 계속 행동합니다.
어쩌면 인간과 알고리즘적 프로필의 가장 중요한 차이가 여기에 있을지도 모릅니다. 프로필은 이미 존재하는 흔적에서 시작해야 합니다. 그러나 인간의 삶은 어제까지 존재하지 않았던 흔적을 여전히 만들어 낼 수 있습니다.
따라서 질문은 단순히 “알고리즘은 나를 아는가?”가 아닙니다.
더 정확한 질문은 이것입니다. “이미 살아온 나로부터 구성된 이야기가, 앞으로 되어 가는 나를 얼마나 담아낼 수 있는가?”
참고문헌
존 체니-리폴드. “새로운 알고리즘적 정체성: 소프트 생명정치와 통제의 조절.” ‘Theory, Culture & Society’, 제28권 제6호, 2011, 164–181쪽. DOI: 10.1177/0263276411424420.
온라인 활동의 패턴으로부터 사용자에 관한 범주를 추론하는 과정과 ‘알고리즘적 정체성’ 개념을 검토하는 데 사용했습니다. 논문 제목은 이해를 돕기 위해 한글로 옮겼습니다.
유럽연합. ‘유럽의회 및 이사회의 규정 (EU) 2016/679’, 2016, 제4조 제4항.
개인의 선호, 관심과 행동 등을 평가·분석·예측하기 위한 자동화된 개인정보 처리를 포함하는 프로파일링의 공식 정의를 확인하는 데 사용했습니다. 자료명은 이해를 돕기 위해 한글로 옮겼습니다.
아얀 신하, 데이비드 F. 글라이히, 카르티크 라마니. “추천 시스템의 피드백 루프 분리.” ‘Advances in Neural Information Processing Systems’, 제29권, 2016, 3243–3251쪽.
추천과 이후의 사용자 행동 사이에서 발생하는 피드백 루프와 추천의 영향을 사용자의 내재적 선호와 구별하는 문제를 검토하는 데 사용했습니다. 논문 제목은 이해를 돕기 위해 한글로 옮겼습니다.
앨리슨 J. B. 체이니, 브랜던 M. 스튜어트, 바버라 E. 엥겔하르트. “추천 시스템의 알고리즘 교란은 어떻게 동질성을 높이고 효용을 낮추는가.” ‘제12회 ACM 추천 시스템 학술대회 논문집’, 2018, 224–232쪽. DOI: 10.1145/3240323.3240370.
이전 추천에 이미 영향을 받은 행동을 다시 학습 데이터로 사용할 때 알고리즘 교란과 행동의 동질화가 발생할 수 있다는 시뮬레이션 결과를 검토하는 데 사용했습니다. 논문 제목은 이해를 돕기 위해 한글로 옮겼습니다.
무즈히 게, 카를라 델가도-바텐펠트, 디트마르 야나흐. “정확성을 넘어서: 범위와 세렌디피티를 통한 추천 시스템 평가.” ‘제4회 ACM 추천 시스템 학술대회 논문집’, 2010, 257–260쪽. DOI: 10.1145/1864708.1864761.
추천 시스템의 품질을 예측 정확도만으로 평가하지 않고 범위와 세렌디피티 같은 차원도 함께 고려할 수 있다는 근거로 사용했습니다. 논문 제목은 이해를 돕기 위해 한글로 옮겼습니다.
Previous Episode - Can Fiction Change Reality? - Stories, Facts, and the Creation of Meaning
Related Label - Narrative
Next Episode -

Comments
Post a Comment