Is My Data Profile the Same as Me?: Behavior, Representation, Inference, and Algorithmic Knowledge
A person stands beyond successive layers of traces, data, models, and predictions, revealing how each representation captures something real while remaining smaller than the life it describes.
Is My Data Profile the Same as Me?: Behavior, Representation, Inference, and Algorithmic Knowledge
Explore how behavior becomes data, data becomes representation, and algorithmic inference turns partial traces into models of a person.
Where Does the Data About Me Begin?
Suppose I search for Nietzsche, watch a lecture, skip another video, buy a novel, listen to a song, and then spend several days reading about an entirely different subject. I experience these actions as parts of a life. A computational system can encounter them differently. Depending on what a service records, some actions may become queries, clicks, viewing times, purchases, ratings, skips, or other interaction data.
Already, a distinction has appeared. The action and the record of the action are not identical.
A search occurred within a situation that included a reason, a moment, competing possibilities, prior knowledge, and a person whose attention could change immediately afterward. The recorded event preserves only what the system was designed and able to record. This does not make the record false. It means that data begins as a representation of something that happened, not as the event in its entirety.
What Was I Able to See Before I Chose?
A second distinction appears before preference can even be inferred. A person can usually choose only among things that somehow become available for attention.
If I did not click on a particular book, several explanations are possible. I may have seen it and rejected it. I may have overlooked it. It may have appeared too far down a list. Or it may never have been shown to me at all. These situations can produce the same observable absence of interaction.
Recommender-system research describes this problem partly through exposure bias. Xu and colleagues emphasize that feedback data in recommender systems are conditioned by what users were exposed to. They show theoretically that supervised learning used to detect user preferences can become inconsistent when the underlying exposure mechanism is ignored.
This means that an unobserved interaction cannot automatically be interpreted as a negative preference. Before asking, “What did the user choose?”, we sometimes have to ask, “What could the user have chosen?” Knowledge about preference therefore requires knowledge about the architecture of exposure.
When Does Behavior Become Data?
Human behavior is continuous, but data is discrete and structured. A person can hesitate, become curious, feel bored, misunderstand an item, return later, or change intention halfway through an action. A system must represent such activity through variables it can process. The particular representation depends on the system.
A star rating explicitly records an evaluation. A click records an interaction but usually not its complete meaning. Viewing duration gives another signal. A purchase provides a different one. A skipped song may indicate dislike, familiarity, distraction, or simply a temporary change in mood.
Recommender-system research therefore distinguishes many forms of explicit and implicit feedback, and different recommendation approaches represent preferences in different ways. Atas and colleagues describe how preferences can be represented as keywords, categories, ratings, attribute values, and critiques depending on the recommendation scenario. They also emphasize that preferences should not always be treated as stable entities waiting to be passively discovered.
This creates another layer in the architecture. The person acts. The system records selected aspects of that action. Those observations are then converted into a form a model can use. None of these stages is identical to the one before it.
Is a Preference Something We Find or Something We Represent?
The word ‘preference’ can make the problem appear simpler than it is. It sounds as though a complete preference already exists somewhere inside a person and the system merely has to discover it. Sometimes this approximation is useful. If I repeatedly choose detective novels over romance novels across many comparable situations, that pattern may be useful for prediction.
But preference representations are still models. Atas and colleagues show that recommender systems can represent preference through substantially different structures depending on how recommendation is performed. They also review psychological research suggesting that preferences can be constructed within decision processes rather than remaining completely stable across contexts. Therefore two systems can possess different representations of the same person without either containing the person in full.
One system may represent me as a vector of item ratings. Another may use categories. Another may infer latent dimensions from patterns of interaction. Another may ask directly what attributes I prefer.
Each representation makes some relationships easier to calculate while leaving others outside the model. Representation is therefore not merely storage. It is a decision about what distinctions will count.
What Does Bias Mean in This Architecture?
Once behavioral data is used to infer preference, the path through which the data was produced becomes important. Chen and colleagues emphasize that recommender-system data is generally observational rather than experimental. Because users do not encounter and evaluate every item under randomly controlled conditions, the observed dataset can contain several forms of bias.
Selection bias can arise because people choose which items to evaluate. Position bias can occur because items shown in prominent positions receive different attention from those shown elsewhere. Exposure bias occurs because users cannot interact with items they never encounter. Popularity bias can amplify already popular items in recommendation data and outputs.
These are not merely defects added after an otherwise perfect representation has been created. They arise partly from the way observations become available in the first place. The architecture of knowledge therefore begins before the model is trained. It begins with the conditions under which the world became observable to the model.
Can Better Models Recover the Person Behind the Data?
A natural response is to build a better model. Recommender-systems research does precisely this. Researchers develop methods to correct bias, estimate exposure, infer counterfactual outcomes, improve preference elicitation, and distinguish signals that ordinary prediction might mix together. Xu and colleagues, for example, use a counterfactual framework designed to account for the exposure mechanism underlying recommendation feedback. Their work demonstrates why ignoring exposure can make preference learning unreliable and proposes a method for addressing the problem.
Such methods can improve inference. But improving inference does not abolish the distinction between model and person. A better map is still a map. A more accurate weather model does not become the atmosphere, and a more accurate recommendation model does not become the human being whose behavior it predicts. The comparison has to be used carefully because a person is not simply another physical system to be modeled in the same way.
The important point is epistemological. Accuracy concerns how well a representation performs with respect to a specified task. Completeness concerns whether the representation contains everything relevant about its object. The first does not automatically establish the second.
What Does a Prediction Actually Tell Us?
Suppose a system predicts with high accuracy that I will click on another article about Nietzsche. What has been established? At minimum, the model has identified information useful for predicting a specified behavior under particular conditions. That can be valuable knowledge. It does not automatically establish why I will click.
Perhaps I admire Nietzsche. Perhaps I disagree with him. Perhaps I am checking a source. Perhaps I am writing critically about the history of philosophy. Perhaps I am tired of seeing Nietzsche recommendations and click only to understand why they keep appearing. Identical actions can arise from different motives. A model does not have to recover every motive in order to predict behavior successfully.
This is why predictive success and explanatory completeness should not be confused. A system can know enough, in a technical sense, to rank an item effectively without possessing anything resembling a complete biography, phenomenology, or philosophical account of the user. In this article, therefore, ‘algorithmic knowledge’ refers only to information generated through data, representation, inference, and prediction. It does not imply human-like understanding.
What Is Lost When One Layer Becomes Another?
Knowledge architecture becomes necessary when different layers are easily mistaken for one another. Consider the sequence again.
A person encounters an environment. Only some possibilities become visible. The person acts. Only some aspects of the action are recorded.
The observations become structured data. The data is represented through selected features or latent variables. A model infers patterns. Those patterns generate predictions. Predictions help rank or recommend future possibilities.
Each transition can preserve useful information. Each transition can also leave something behind.
The environment contains possibilities the person never encountered. The person's action contains motives the log never records. The data contains patterns the model may fail to represent. The model contains abstractions that do not exist inside the person in the same form. The prediction describes a probability or ranking, not a completed future.
The architecture is therefore not a chain of identical copies. It is a chain of transformations.
Can an Algorithmic Profile Still Be Useful?
Recognizing these limits does not make algorithmic profiles worthless. A map becomes useful precisely because it does not reproduce every stone, tree, sound, and movement in the territory. It reduces. Good models also reduce. The question is whether the reduction preserves what matters for the task for which the model is being used.
If a system recommends music, listening history may be useful. If it recommends medical treatment, music-listening behavior may be irrelevant while other kinds of evidence become essential. Knowledge must therefore be evaluated relative to its object, purpose, evidence, and method.
Problems arise when a representation built for one purpose silently acquires a larger meaning. A profile designed to predict clicks can begin to be treated as a description of interest.
Interest can then be treated as preference. Preference can be treated as identity. Identity can finally be treated as the person. Every step may sound plausible.
But each step requires a new justification. Without that justification, the architecture collapses distinct levels into one another.
Where Does the Human Remain Outside the Model?
There is another reason a user profile cannot be treated as a completed person. The person continues to act after the profile is constructed.
People encounter new circumstances. They learn. They forget. They deliberately change habits. They discover interests that were absent from previous data. They reject things they once enjoyed. They sometimes act inconsistently even by their own standards.
A model can be updated when new traces appear. But updating does not remove the temporal gap completely. The model is always constructed from information that has become available in some form. The next action has not yet become data.
This is not a defect unique to algorithms. Every description of a living person faces a similar problem.
A diary ends at the last written sentence. A biography stops at its date of completion. A psychological measurement describes a person under particular conditions. A statistical model contains what its variables and observations allow it to contain. A living person continues beyond all of them.
How Should Knowledge About People Be Organized?
The answer is not to reject data. It is to preserve distinctions.
Behavior is not identical to preference. Non-interaction is not necessarily rejection. Observed preference is not independent of exposure. A data profile is not identical to a person.
A model is not identical to its object. Prediction is not identical to explanation. Accuracy is not identical to completeness. And computational inference is not identical to human understanding.
Once these boundaries are preserved, different kinds of knowledge can be connected without being confused. Behavioral data can tell us something real. Exposure data can tell us what opportunities preceded the behavior. Preference representations can organize recurring patterns. Causal methods can help separate relationships that observational data alone cannot establish. Predictive models can estimate what may happen next. Human interpretation can ask what those patterns mean within a life.
No single layer has to perform every task. That is precisely why architecture is necessary.
Is My Data Profile the Same as Me?
The answer is no, but that answer is only the beginning. A useful profile does not have to be identical to a person. The more important question is what claims the profile legitimately supports.
If a model predicts that I will read another article about Nietzsche, it may have learned something useful about my observable behavior. It has not thereby established the complete meaning of Nietzsche in my life. Nor has it established that the preference will remain tomorrow. It has created a representation from available evidence for a particular inferential purpose.
That representation can influence what I encounter next, as the earlier articles in this series have shown. For that reason, its limits matter as much as its accuracy. Knowledge architecture begins when we stop asking one representation to stand in for an entire reality.
The final question is therefore not merely, “How accurately can an algorithm model me?” It is, “What has been preserved, what has been inferred, and what has disappeared each time a human life is translated into data?”
References
Atas, Müslüm, Alexander Felfernig, Seda Polat-Erdeniz, Andrei Popescu, Thi Ngoc Trang Tran, and Mathias Uta. “Towards Psychology-Aware Preference Construction in Recommender Systems: Overview and Research Issues.” ‘Journal of Intelligent Information Systems’, vol. 57, no. 3, 2021, pp. 467–489. DOI: 10.1007/s10844-021-00674-5.
This article was used to distinguish users from formal preference representations and to verify that recommender systems represent preferences through different structures, including keywords, categories, ratings, attribute values, and critiques. It was also used to examine preference construction rather than assuming completely fixed preferences.
Chen, Jiawei, Hande Dong, Xiang Wang, Fuli Feng, Meng Wang, and Xiangnan He. “Bias and Debias in Recommender System: A Survey and Future Directions.” ‘ACM Transactions on Information Systems’, vol. 41, no. 3, 2023, Article 67, pp. 1–39. DOI: 10.1145/3564284.
This article was used to establish that recommender-system behavior data is observational rather than experimental and to distinguish selection, exposure, position, popularity, and other forms of bias that can affect inference from user behavior.
Xu, Da, Chuanwei Ruan, Evren Korpeoglu, Sushant Kumar, and Kannan Achan. “Adversarial Counterfactual Learning and Evaluation for Recommender System.” ‘Advances in Neural Information Processing Systems’, vol. 33, 2020, pp. 13515–13526.
This study was used to verify that recommendation feedback is conditioned by exposure, that preference learning can become inconsistent when exposure information is ignored, and that counterfactual methods can be used to model the underlying exposure mechanism.
한 사람이 흔적, 데이터, 모형과 예측의 연속된 층위 너머에 서 있으며, 각각의 표상이 실제의 일부를 포착하면서도 그것이 설명하는 삶 전체보다는 작다는 사실을 보여 줍니다.
나의 데이터 프로필은 나와 같은가? - 행동, 표상, 추론, 알고리즘적 지식
행동이 데이터가 되고, 데이터가 표상이 되며, 알고리즘적 추론이 부분적인 흔적을 인간에 관한 모형으로 바꾸는 과정을 살펴봅니다.
나에 관한 데이터는 어디에서 시작될까?
제가 니체를 검색하고, 강의를 하나 보고, 다른 영상은 건너뛰고, 소설을 한 권 구매한 뒤, 음악을 듣고 며칠 동안 전혀 다른 주제를 읽었다고 가정해 보겠습니다. 저는 이러한 행동을 하나의 삶을 이루는 부분으로 경험합니다. 컴퓨터 시스템은 그것을 다른 방식으로 접할 수 있습니다. 서비스가 무엇을 기록하도록 설계되었는지에 따라 일부 행동은 검색어, 클릭, 시청 시간, 구매, 평점, 건너뛰기 또는 다른 상호작용 데이터가 될 수 있습니다.
벌써 하나의 구분이 나타납니다. 행동과 행동의 기록은 동일하지 않습니다.
검색은 하나의 상황 안에서 이루어졌습니다. 거기에는 이유와 순간, 경쟁하는 다른 가능성, 기존 지식과 바로 다음 순간에 관심을 바꿀 수 있는 인간이 있었습니다. 기록된 사건은 시스템이 기록하도록 설계되었고 실제로 기록할 수 있었던 것만을 보존합니다. 그렇다고 기록이 거짓이라는 뜻은 아닙니다. 데이터는 일어난 사건 전체가 아니라 그 사건을 표상한 것으로 시작된다는 뜻입니다.
나는 선택하기 전에 무엇을 볼 수 있었을까?
선호를 추론하기 전에도 두 번째 구분이 나타납니다. 사람은 보통 어떤 방식으로든 자신의 주의에 들어온 것 가운데 선택할 수 있습니다.
제가 어떤 책을 클릭하지 않았다면 여러 설명이 가능합니다. 그 책을 보고 거부했을 수도 있고, 그냥 지나쳤을 수도 있습니다. 목록의 너무 아래쪽에 나타났을 수도 있으며, 애초에 저에게 전혀 제시되지 않았을 수도 있습니다. 이러한 상황들은 모두 ‘상호작용이 없었다’는 동일한 관찰 결과를 만들 수 있습니다.
추천 시스템 연구에서는 이 문제를 부분적으로 ‘노출 편향’이라는 개념으로 다룹니다. 쉬와 동료들은 추천 시스템의 피드백 데이터가 사용자가 실제로 무엇에 노출되었는지에 영향을 받는다는 점을 강조합니다. 이들은 기저에 있는 노출 메커니즘을 고려하지 않을 경우 사용자 선호를 탐지하기 위한 지도학습에서 일관되지 않은 결과가 나타날 수 있음을 이론적으로 보였습니다.
따라서 관찰되지 않은 상호작용을 자동으로 부정적인 선호라고 해석할 수 없습니다. “사용자는 무엇을 선택했는가?”라고 묻기 전에 때로는 “사용자는 무엇을 선택할 수 있었는가?”라고 물어야 합니다. 그러므로 선호에 관한 지식에는 노출 구조에 관한 지식이 필요합니다.
행동은 언제 데이터가 될까?
인간의 행동은 연속적이지만 데이터는 분리되고 구조화됩니다. 사람은 망설이거나, 호기심을 느끼거나, 지루해지거나, 대상을 잘못 이해하거나, 나중에 다시 돌아오거나, 행동 도중에 의도를 바꿀 수 있습니다. 시스템은 이러한 활동을 자신이 처리할 수 있는 변수로 표현해야 합니다. 구체적인 표상 방식은 시스템에 따라 달라집니다.
별점은 평가를 명시적으로 기록합니다. 클릭은 상호작용을 기록하지만 대개 그 의미 전체를 기록하지는 않습니다. 시청 시간은 또 다른 신호를 제공합니다. 구매는 다른 종류의 신호입니다. 음악을 건너뛴 행동은 싫어한다는 뜻일 수도 있지만, 이미 잘 알고 있거나, 다른 데 정신이 팔렸거나, 잠시 기분이 달라졌다는 뜻일 수도 있습니다.
따라서 추천 시스템 연구에서는 다양한 명시적·암묵적 피드백을 구분하며, 추천 방식에 따라 선호를 서로 다른 방식으로 표현합니다. 아타스와 동료들은 추천 상황에 따라 선호가 키워드, 범주, 평점, 속성값과 비판적 피드백 등으로 표현될 수 있다고 설명합니다. 이들은 또한 선호를 수동적으로 발견되기를 기다리는 안정된 실체로만 취급해서는 안 된다는 점을 강조합니다.
여기에서 구조의 또 다른 층위가 나타납니다. 인간이 행동합니다. 시스템은 그 행동 가운데 선택된 일부를 기록합니다. 그 관찰값은 다시 모형이 사용할 수 있는 형태로 변환됩니다. 이 단계 가운데 어느 것도 바로 앞 단계와 동일하지 않습니다.
선호는 발견되는 것일까, 표상되는 것일까?
‘선호’라는 단어는 문제를 실제보다 단순하게 보이게 할 수 있습니다. 완전한 선호가 이미 인간 내부 어딘가에 존재하고 시스템은 그것을 발견하기만 하면 되는 것처럼 들리기 때문입니다. 때로는 이러한 근사가 유용합니다. 제가 비교 가능한 여러 상황에서 로맨스 소설보다 탐정소설을 반복적으로 선택한다면 그 패턴은 예측에 유용할 수 있습니다.
그러나 선호의 표상은 여전히 모형입니다. 아타스와 동료들은 추천 방식에 따라 추천 시스템이 매우 다른 구조로 선호를 표현할 수 있음을 보여 줍니다. 또한 심리학 연구를 검토하면서 선호가 모든 맥락에서 완전히 안정된 상태로 유지되기보다는 의사결정 과정에서 구성될 수 있다고 설명합니다. 따라서 두 시스템이 동일한 사람에 관해 서로 다른 표상을 가지고 있으면서도 어느 쪽도 그 인간 전체를 담고 있지 않을 수 있습니다.
한 시스템은 저를 항목별 평점 벡터로 표현할 수 있습니다. 다른 시스템은 범주를 사용할 수 있습니다. 또 다른 시스템은 상호작용 패턴에서 잠재적인 차원을 추론할 수 있습니다. 다른 시스템은 제가 어떤 속성을 선호하는지 직접 질문할 수도 있습니다.
각각의 표상은 특정한 관계를 더 쉽게 계산할 수 있게 하지만 다른 관계를 모형 밖에 남깁니다. 따라서 표상은 단순한 저장이 아닙니다. 어떤 차이를 의미 있는 것으로 취급할 것인지를 결정하는 일이기도 합니다.
이 구조에서 편향은 무엇을 의미할까?
행동 데이터에서 선호를 추론하기 시작하면 그 데이터가 어떤 경로로 만들어졌는지가 중요해집니다. 첸과 동료들은 추천 시스템 데이터가 일반적으로 실험 데이터가 아니라 관찰 데이터라는 사실을 강조합니다. 사용자가 무작위로 통제된 조건에서 모든 항목을 접하고 평가하는 것이 아니기 때문에 관찰 데이터에는 여러 형태의 편향이 존재할 수 있습니다.
선택 편향은 사람들이 어떤 항목을 평가할 것인지 스스로 선택하기 때문에 발생할 수 있습니다. 위치 편향은 눈에 잘 띄는 위치에 제시된 항목이 다른 위치의 항목과 다른 정도의 주의를 받기 때문에 나타날 수 있습니다. 노출 편향은 사용자가 전혀 접하지 않은 항목과 상호작용할 수 없기 때문에 발생합니다. 인기 편향은 이미 인기 있는 항목이 추천 데이터와 결과에서 더욱 강화되는 현상과 관련됩니다.
이러한 문제는 완벽한 표상이 만들어진 뒤 추가되는 단순한 결함이 아닙니다. 무엇이 관찰될 수 있었는지를 결정하는 과정 자체에서 부분적으로 발생합니다. 따라서 지식의 구조는 모형이 학습되기 전부터 시작됩니다. 세계가 어떤 조건에서 모형에게 관찰될 수 있었는지에서 시작됩니다.
더 좋은 모형은 데이터 뒤의 인간을 복원할 수 있을까?
자연스러운 대응은 더 좋은 모형을 만드는 것입니다. 추천 시스템 연구도 바로 그렇게 합니다. 연구자들은 편향을 교정하고, 노출을 추정하고, 반사실적 결과를 추론하고, 선호를 더 정확하게 파악하며, 일반적인 예측에서는 섞일 수 있는 신호들을 분리하는 방법을 개발합니다. 예를 들어 쉬와 동료들은 추천 피드백의 기저에 존재하는 노출 메커니즘을 고려하도록 설계된 반사실적 방법을 사용합니다. 이들의 연구는 노출을 무시할 때 선호 학습이 왜 불안정해질 수 있는지를 보여 주고, 이를 다루기 위한 방법을 제시합니다.
이러한 방법은 추론을 개선할 수 있습니다. 그러나 추론의 개선이 모형과 인간의 구별을 없애는 것은 아닙니다. 더 정확한 지도도 여전히 지도입니다. 더 정확한 기상 모형이 대기 자체가 되지 않는 것처럼, 더 정확한 추천 모형도 자신이 행동을 예측하는 인간 자체가 되지는 않습니다. 다만 이 비교 역시 조심해서 사용해야 합니다. 인간은 단순히 같은 방식으로 모형화되는 또 하나의 물리계가 아니기 때문입니다.
중요한 것은 인식론적인 차이입니다. 정확성은 특정한 과업에 대해 표상이 얼마나 잘 작동하는지와 관련됩니다. 완전성은 그 표상이 대상에 관한 모든 관련 요소를 포함하고 있는지와 관련됩니다. 첫 번째가 성립한다고 두 번째까지 자동으로 성립하는 것은 아닙니다.
예측은 실제로 무엇을 알려 줄까?
한 시스템이 제가 니체에 관한 다음 글을 클릭할 것이라고 높은 정확도로 예측한다고 가정해 보겠습니다. 무엇이 확인된 것일까요? 최소한 그 모형은 특정한 조건에서 특정 행동을 예측하는 데 유용한 정보를 발견했습니다. 그것은 가치 있는 지식일 수 있습니다. 그러나 그것만으로 제가 왜 클릭할지는 확정되지 않습니다.
제가 니체를 좋아할 수도 있습니다. 반대할 수도 있습니다. 출처를 확인하고 있을 수도 있습니다. 철학사를 비판적으로 쓰고 있을 수도 있습니다. 계속 니체가 추천되는 것이 지겨워서 왜 반복되는지 확인하려고 클릭할 수도 있습니다. 같은 행동은 서로 다른 동기에서 나올 수 있습니다. 모형이 행동을 성공적으로 예측하기 위해 모든 동기를 복원해야 하는 것은 아닙니다.
그러므로 예측의 성공과 설명의 완전성을 혼동해서는 안 됩니다. 시스템은 기술적인 의미에서 어떤 항목의 순위를 효과적으로 결정할 만큼 충분한 것을 파악할 수 있지만, 그 사용자에 관한 완전한 전기나 경험의 구조 또는 철학적 설명을 가지고 있지 않을 수 있습니다. 따라서 이 글에서 ‘알고리즘적 지식’은 데이터, 표상, 추론과 예측을 통해 생성된 정보만을 의미합니다. 인간과 같은 이해를 의미하지 않습니다.
하나의 층위가 다른 층위로 바뀔 때 무엇이 사라질까?
서로 다른 층위를 같은 것으로 착각하기 쉬울 때 지식 아키텍처가 필요합니다. 앞의 순서를 다시 생각해 보겠습니다.
인간은 하나의 환경을 접합니다. 그 가운데 일부 가능성만 보입니다. 인간이 행동합니다. 그 행동 가운데 일부만 기록됩니다.
관찰값이 구조화된 데이터가 됩니다. 데이터는 선택된 특징이나 잠재 변수로 표현됩니다. 모형이 패턴을 추론합니다. 그 패턴으로 예측이 만들어집니다. 예측은 다음에 어떤 가능성을 높은 순위로 보여 주거나 추천할지에 관여합니다.
각각의 변환은 유용한 정보를 보존할 수 있습니다. 각 변환에서 무엇인가가 남겨질 수도 있습니다.
환경에는 인간이 전혀 접하지 못한 가능성이 있습니다. 인간의 행동에는 기록에 들어가지 않은 동기가 있습니다. 데이터에는 모형이 표현하지 못한 패턴이 있을 수 있습니다. 모형에는 인간 내부에 동일한 형태로 존재하지 않는 추상적 구조가 포함될 수 있습니다. 예측은 확률이나 순위를 나타내는 것이지 이미 완성된 미래를 나타내는 것은 아닙니다.
따라서 이 구조는 동일한 복사본이 이어지는 사슬이 아닙니다. 서로 다른 변환이 이어지는 사슬입니다.
그렇다면 알고리즘적 프로필은 여전히 유용할까?
이러한 한계를 인정한다고 알고리즘적 프로필이 무용해지는 것은 아닙니다. 지도는 실제 영토의 모든 돌과 나무, 소리와 움직임을 재현하지 않기 때문에 오히려 유용해집니다. 좋은 모형도 축약합니다. 문제는 그 축약이 모형이 사용되는 과업에 필요한 것을 제대로 보존하는가입니다.
음악을 추천하는 시스템이라면 청취 기록이 유용할 수 있습니다. 의학적 치료를 추천한다면 음악 청취 기록은 관련이 없을 수 있으며 전혀 다른 종류의 근거가 중요해집니다. 따라서 지식은 대상, 목적, 근거와 방법에 맞추어 평가해야 합니다.
문제는 특정한 목적을 위해 만들어진 표상이 조용히 더 큰 의미를 갖기 시작할 때 생깁니다. 클릭을 예측하기 위해 만들어진 프로필이 관심에 관한 설명으로 취급될 수 있습니다.
관심이 다시 취향으로 취급될 수 있습니다. 취향이 정체성으로 취급될 수 있습니다. 마지막에는 정체성이 인간 그 자체로 취급될 수 있습니다. 각 단계는 그럴듯하게 들릴 수 있습니다.
그러나 각각의 단계에는 새로운 정당화가 필요합니다. 그 정당화가 없다면 서로 다른 층위가 하나로 붕괴합니다.
모형 밖에는 인간의 무엇이 남을까?
사용자 프로필을 완성된 인간으로 취급할 수 없는 또 하나의 이유가 있습니다. 프로필이 만들어진 뒤에도 인간은 계속 행동합니다.
사람은 새로운 환경을 만납니다. 배웁니다. 잊습니다. 의도적으로 습관을 바꿉니다. 과거 데이터에는 존재하지 않았던 관심사를 발견합니다. 예전에 좋아했던 것을 거부하기도 합니다. 때로는 자신의 기준으로 보아도 일관되지 않게 행동합니다.
새로운 흔적이 생기면 모형도 업데이트할 수 있습니다. 그러나 업데이트한다고 시간의 간격이 완전히 사라지는 것은 아닙니다. 모형은 어떤 형태로든 이미 이용 가능해진 정보에서 만들어지기 때문입니다. 다음 행동은 아직 데이터가 되지 않았습니다.
이것은 알고리즘에만 존재하는 결함도 아닙니다. 살아 있는 인간을 설명하는 모든 방식에는 비슷한 문제가 있습니다.
일기는 마지막으로 적은 문장에서 끝납니다. 전기는 완성된 시점에서 멈춥니다. 심리 측정은 특정한 조건에서 인간을 기술합니다. 통계적 모형은 자신의 변수와 관찰값이 허용하는 것을 담습니다. 살아 있는 인간은 그 모든 것의 다음으로 계속 나아갑니다.
인간에 관한 지식은 어떻게 구조화해야 할까?
해답은 데이터를 거부하는 것이 아닙니다. 구별을 유지하는 것입니다.
행동은 선호와 동일하지 않습니다. 상호작용이 없다는 것이 반드시 거부를 뜻하지는 않습니다. 관찰된 선호는 노출과 독립되어 있지 않을 수 있습니다. 데이터 프로필은 인간과 동일하지 않습니다.
모형은 자신의 대상과 동일하지 않습니다. 예측은 설명과 동일하지 않습니다. 정확성은 완전성과 동일하지 않습니다. 컴퓨터의 추론은 인간의 이해와 동일하지 않습니다.
이러한 경계를 유지하면 서로 다른 종류의 지식을 혼동하지 않고 연결할 수 있습니다. 행동 데이터는 실제의 일부를 알려 줄 수 있습니다. 노출 데이터는 그 행동에 앞서 어떤 기회가 존재했는지 알려 줄 수 있습니다. 선호 표상은 반복되는 패턴을 조직할 수 있습니다. 인과적 방법은 관찰 자료만으로는 확정하기 어려운 관계를 구분하는 데 도움을 줄 수 있습니다. 예측 모형은 다음에 무엇이 일어날 가능성이 있는지 추정할 수 있습니다. 인간의 해석은 그러한 패턴이 하나의 삶에서 무엇을 의미하는지 질문할 수 있습니다.
하나의 층위가 모든 일을 수행할 필요는 없습니다. 바로 그렇기 때문에 구조가 필요합니다.
나의 데이터 프로필은 나와 같은가?
답은 ‘아니다’입니다. 그러나 그 답은 시작일 뿐입니다. 유용한 프로필이 인간과 동일할 필요는 없습니다. 더 중요한 문제는 그 프로필로부터 어떤 주장까지 정당하게 할 수 있는가입니다.
모형이 제가 니체에 관한 다음 글을 읽을 것이라고 예측했다면, 저의 관찰 가능한 행동에 관해 유용한 무엇인가를 학습했을 수 있습니다. 그렇다고 제 삶에서 니체가 지닌 의미 전체까지 확립한 것은 아닙니다. 그 취향이 내일까지 계속될 것이라는 사실도 확립하지 못했습니다. 이용 가능한 근거로부터 특정한 추론 목적을 위한 표상을 만든 것입니다.
앞의 글들에서 살펴보았듯이 그러한 표상은 제가 다음에 무엇을 접하는지에도 관여할 수 있습니다. 그렇기 때문에 정확성만큼 그 한계도 중요합니다. 하나의 표상에게 현실 전체를 대신하게 하지 않을 때 지식 아키텍처가 시작됩니다.
따라서 마지막 질문은 단지 “알고리즘은 나를 얼마나 정확하게 모형화할 수 있는가?”가 아닙니다. 더 중요한 질문은 이것입니다. “하나의 인간의 삶이 데이터로 번역될 때마다 무엇이 보존되고, 무엇이 추론되며, 무엇이 사라지는가?”

Comments
Post a Comment