The Reality of Data: The Boundary Between Factual Truth and Measurable Interpretation
A measurement philosophy infographic showing the process of nature's reality being converted into data and lost through the grid of measurement tools.
The Reality of Data: The Boundary Between Factual Truth and Measurable Interpretation
Data is a measured representation of reality, shaped by definitions, sampling, instruments, and analytical methods.
Tracing the Origins: Phenomenological Observation and the Construction of Data
The proposition that data is not reality itself but the result of observation or measurement lies at the intersection of realism and constructivism in the philosophy of science. According to Edmund Husserl’s phenomenology, consciousness is always directed toward something, and objects are experienced through intentional structures rather than received as uninterpreted contents.
This does not mean that consciousness freely invents reality. It means that human beings encounter reality through particular forms of attention, perception, and interpretation. The further claim that instruments mediate how the world becomes observable belongs more directly to the philosophy of technology. A measuring instrument does not reproduce the whole of a phenomenon. It makes particular properties observable by translating them into recorded values.
Physical measurement also reveals that observation has limits. In quantum mechanics, Heisenberg’s uncertainty principle establishes an intrinsic limit on the simultaneous precision with which certain pairs of observables, such as position and momentum, can be specified. Quantum measurement may also alter the state being measured, but this measurement effect should not be treated as identical to the uncertainty principle.
The uncertainty principle does not prove that all data is subjective or that reality is created by observation. It demonstrates that even rigorous physical measurement operates within conditions that determine what can be specified and with what degree of precision. Data therefore preserves neither the totality of an event nor reality in an unmediated form. It records particular properties under defined conditions.
The philosophy of statistician W. Edwards Deming supports a related conclusion. Deming emphasized the importance of operational definitions: agreed procedures that translate concepts into measurable forms. Before a value can be interpreted, one must understand what was measured, how it was measured, and what process produced the result.
Data should therefore be understood as a product of interaction among a phenomenon, a method, an instrument, and an interpretive framework. The relevant question is not merely whether data is authentic, but through which measurement logic it was produced.
Genealogical Contemplation: The Historical Authority of Objectivity
Throughout modern history, data has acquired authority because it appears to provide objective descriptions of reality. As bureaucracies, scientific institutions, businesses, and governments increasingly organized society through statistics, measurement became an essential instrument of administration and decision-making.
Quantification made complex societies more manageable. Populations could be counted, resources compared, diseases tracked, production evaluated, and economic changes represented across time. Yet every system of measurement required categories. Before people could be counted, institutions had to decide which distinctions mattered and how individuals would be classified.
The history of data therefore developed alongside the history of instruments and institutions. Rulers, scales, clocks, censuses, surveys, sensors, and computational systems expanded the range and precision of what could be recorded. Greater precision improved knowledge, but it did not eliminate decisions about what should be measured or how results should be interpreted.
Large language models belong to this history only in a qualified sense. They are not measuring instruments in the same way as rulers, scales, clocks, or sensors. They are computational systems that analyze, transform, and generate information from previously produced data. When connected to observational systems, they may participate in a measurement process, but they do not independently measure reality merely by processing language.
Structural distortion occurs when selected observations are treated as complete representations of a phenomenon. What remains unmeasured may disappear from administrative attention even though it continues to exist in reality. Qualities that resist quantification can consequently receive less recognition than variables that are easily recorded and compared.
Data does not speak independently. Its meaning depends upon the systems that define categories, collect observations, process values, and determine which results deserve attention. The history of data is therefore also a history of the changing boundaries of institutional and technological visibility.
Declaration of Interpretation: Measurement and Inevitable Exclusion
Recognizing data as the result of measurement does not require distrusting it. It requires asking more precise questions about how it was produced and what its evidential limits are.
Every dataset results from decisions about what to observe, how to measure it, which population or events to include, and what to leave unrecorded. Selection and omission are therefore unavoidable features of data production. This does not make data false. It establishes the range within which its conclusions remain justified.
A temperature measurement does not describe every physical property of an object. An unemployment rate does not represent every dimension of economic insecurity. A standardized test does not capture the whole of a person’s intelligence. Each measure isolates particular features because no single dataset can preserve the totality of a phenomenon.
Interpreting data therefore requires reconstructing its measurement design. We must examine its operational definitions, sampling procedures, instruments, units, uncertainty, missing values, processing methods, and analytical assumptions. Without this information, numerical precision can produce confidence unsupported by the measurement process.
The danger begins when a partial representation is mistaken for a complete reality. The more heavily a society relies on quantified information, the more carefully it must distinguish what the data directly supports from what has been inferred beyond it.
Data is a tool, and every tool makes some aspects of the world easier to perceive while leaving others less visible. Measurement expands human knowledge, but it also establishes a boundary around what a particular method can reveal.
Modern Application: Data Literacy in the Age of Information Overload
Modern societies receive continuous streams of numerical, textual, visual, and behavioral data. Governments, businesses, scientists, and artificial intelligence systems use this material to identify patterns, compare alternatives, and make predictions. Data-driven analysis can improve decisions, but its outputs remain dependent upon the conditions under which the underlying data was produced.
What, therefore, must we do? First, we must examine not only the source of data but also its methodology. Second, we must identify the populations, variables, experiences, and uncertainties that the data failed to capture. Third, we must treat data-driven conclusions as claims supported within defined limits rather than as absolute truths.
This attitude does not amount to doubting every number. It distinguishes responsible interpretation from passive submission to numerical authority. A person who treats data as reality may become governed by categories they have never examined. A person who understands data as a measured representation can evaluate its limits, compare it with other forms of evidence, and decide how much authority it deserves.
Data is not a perfect mirror of reality. It is a lens constructed through observation, measurement, classification, and interpretation. Some lenses are more accurate and reliable than others, but none reveals every aspect of the world at once.
Only by understanding the curvature of the lens can we form a more warranted interpretation of the reality that the data partially represents. As long as we continue to question how data is produced, it can remain a tool that supports thought rather than an authority that replaces it.
References:
Husserl, E. (1982). 'Ideas Pertaining to a Pure Phenomenology and to a Phenomenological Philosophy, First Book.' Translated by F. Kersten. Martinus Nijhoff Publishers.
Ihde, D. (1990). 'Technology and the Lifeworld: From Garden to Earth.' Indiana University Press.
Heisenberg, W. (1958). 'Physics and Philosophy: The Revolution in Modern Science.' Harper & Brothers.
Deming, W. E. (1994). 'The New Economics for Industry, Government, Education.' Second edition. MIT Press.
Porter, T. M. (1995). 'Trust in Numbers: The Pursuit of Objectivity in Science and Public Life.' Princeton University Press.
Bowker, G. C., & Star, S. L. (1999). 'Sorting Things Out: Classification and Its Consequences.' MIT Press.
Based on Edmund Husserl’s analysis of intentionality, Don Ihde’s philosophy of technological mediation, Werner Heisenberg’s account of the limits of quantum measurement, W. Edwards Deming’s theory of operational definitions and measurement systems, Theodore Porter’s history of quantitative objectivity, and Geoffrey Bowker and Susan Leigh Star’s analysis of classification systems. The discussion of large language models, data literacy, and contemporary data-driven decision-making is an extension of these frameworks and is not directly derived from these works.
자연의 실재가 데이터로 변환되는 과정과, 그 과정에서 측정 도구의 격자(grid)를 통해 정보가 소실되는 모습을 보여주는 측정 철학 인포그래픽.
데이터의 실재: 사실적 진실과 측정 가능한 해석의 경계
데이터는 정의, 표본 선정, 도구, 분석 방법에 의해 형성된 현실의 측정 표상입니다.
원전의 추적: 현상학적 관찰과 데이터의 구성
데이터가 현실 자체가 아니라 관찰이나 측정의 결과라는 명제는 과학철학에서 실재론과 구성주의가 교차하는 지점에 놓여 있습니다. 에드문트 후설의 현상학에 따르면 의식은 언제나 무엇인가를 향하며, 대상은 해석되지 않은 내용으로 그대로 받아들여지는 것이 아니라 지향적 구조를 통해 경험됩니다.
이것은 의식이 현실을 자유롭게 만들어 낸다는 뜻이 아닙니다. 인간이 특정한 주의와 지각, 해석의 형식을 통해 현실을 만난다는 뜻입니다. 도구가 세계를 관찰 가능한 형태로 매개한다는 주장은 기술철학에 더 직접적으로 속합니다. 측정 도구는 현상의 전체를 재현하지 않습니다. 특정한 속성을 기록된 값으로 변환하여 관찰할 수 있게 합니다.
물리적 측정 역시 관찰에 한계가 있음을 보여줍니다. 양자역학에서 하이젠베르크의 불확정성 원리는 위치와 운동량처럼 특정한 관측 가능량의 쌍을 동시에 확정할 수 있는 정밀도에 본질적인 한계가 있음을 규정합니다. 양자 측정은 측정되는 상태를 변화시킬 수도 있지만, 이러한 측정 효과를 불확정성 원리와 동일한 것으로 취급해서는 안 됩니다.
불확정성 원리는 모든 데이터가 주관적이거나 현실이 관찰을 통해 만들어진다는 것을 증명하지 않습니다. 그것은 엄밀한 물리적 측정조차 무엇을 어느 정도의 정밀도로 확정할 수 있는지를 결정하는 조건 안에서 작동한다는 사실을 보여줍니다. 따라서 데이터는 사건의 전체나 매개되지 않은 형태의 현실을 보존하지 않습니다. 데이터는 규정된 조건에서 특정한 속성을 기록합니다.
통계학자 W. 에드워즈 데밍의 철학도 이와 관련된 결론을 뒷받침합니다. 데밍은 개념을 측정 가능한 형태로 변환하는 합의된 절차인 조작적 정의의 중요성을 강조했습니다. 하나의 값을 해석하기 전에 무엇을 측정했고, 어떻게 측정했으며, 어떤 과정이 그 결과를 만들어 냈는지 이해해야 합니다.
따라서 데이터는 현상·방법·도구·해석 체계 사이의 상호작용이 만들어 낸 산물로 이해해야 합니다. 중요한 질문은 데이터가 진짜인지에만 있지 않고, 어떤 측정 논리를 통해 생산되었는지에도 있습니다.
계보적 고찰: 객관성이 지닌 역사적 권위
근대 역사에서 데이터는 현실에 대한 객관적 설명을 제공하는 것처럼 보였기 때문에 권위를 획득했습니다. 관료제와 과학기관, 기업과 정부가 통계를 통해 사회를 조직하는 일이 늘어나면서 측정은 행정과 의사결정의 핵심 수단이 되었습니다.
정량화는 복잡한 사회를 더욱 관리하기 쉽게 만들었습니다. 인구를 계산하고, 자원을 비교하고, 질병을 추적하고, 생산을 평가하며, 경제적 변화를 시간에 따라 나타낼 수 있게 되었습니다. 그러나 모든 측정 체계에는 범주가 필요했습니다. 사람을 계산하기 전에 기관은 어떤 차이가 중요하며 개인을 어떻게 분류할 것인지 결정해야 했습니다.
따라서 데이터의 역사는 도구와 제도의 역사와 함께 발전했습니다. 자와 저울, 시계와 인구조사, 설문조사와 센서, 계산 시스템은 기록할 수 있는 대상의 범위와 정밀도를 확장했습니다. 정밀도의 향상은 지식을 개선했지만, 무엇을 측정하고 결과를 어떻게 해석할 것인지에 관한 결정을 제거하지는 못했습니다.
대규모 언어 모델은 제한된 의미에서만 이 역사에 포함됩니다. 대규모 언어 모델은 자·저울·시계·센서와 같은 방식으로 작동하는 측정 도구가 아닙니다. 이미 생산된 데이터에서 정보를 분석하고 변환하며 생성하는 계산 시스템입니다. 관찰 시스템과 연결될 경우 측정 과정에 참여할 수 있지만, 언어를 처리한다는 이유만으로 현실을 독립적으로 측정하는 것은 아닙니다.
선택된 관찰 결과를 하나의 현상에 대한 완전한 표상으로 취급할 때 구조적 왜곡이 발생합니다. 측정되지 않은 것은 현실에 계속 존재하더라도 행정적 관심에서 사라질 수 있습니다. 그 결과 정량화하기 어려운 특성은 쉽게 기록하고 비교할 수 있는 변수보다 적은 관심을 받을 수 있습니다.
데이터는 독립적으로 말하지 않습니다. 데이터의 의미는 범주를 정의하고, 관찰 결과를 수집하고, 값을 처리하며, 어떤 결과에 주목할 것인지를 결정하는 체계에 의존합니다. 따라서 데이터의 역사는 제도적·기술적 가시성의 경계가 변화해 온 역사이기도 합니다.
해석 선언: 측정과 불가피한 배제
데이터를 측정의 결과로 인식한다고 해서 데이터를 불신해야 하는 것은 아닙니다. 데이터가 어떻게 생산되었으며 그 증거가 어디까지 유효한지에 관해 더욱 정확한 질문을 해야 한다는 뜻입니다.
모든 데이터세트는 무엇을 관찰하고, 어떻게 측정하고, 어떤 모집단이나 사건을 포함하며, 무엇을 기록하지 않을 것인지에 관한 결정에서 만들어집니다. 따라서 선택과 누락은 데이터 생산에서 피할 수 없는 특성입니다. 그렇다고 데이터가 거짓이 되는 것은 아닙니다. 이는 데이터에서 도출한 결론이 정당화될 수 있는 범위를 규정합니다.
온도 측정은 한 물체의 모든 물리적 특성을 설명하지 않습니다. 실업률은 경제적 불안정의 모든 차원을 나타내지 않습니다. 표준화된 검사는 한 사람의 지능 전체를 포착하지 않습니다. 하나의 데이터세트가 현상의 전체를 보존할 수 없기 때문에 각각의 측정은 특정한 특성만을 분리합니다.
따라서 데이터를 해석하려면 그 측정 설계를 재구성해야 합니다. 조작적 정의와 표본 추출 절차, 측정 도구와 단위, 불확실성과 결측값, 처리 방법과 분석 가정을 검토해야 합니다. 이러한 정보가 없으면 수치의 정밀성이 측정 과정으로 뒷받침되지 않는 확신을 만들어 낼 수 있습니다.
부분적인 표상을 완전한 현실로 오인하는 순간 위험이 시작됩니다. 사회가 정량화된 정보에 크게 의존할수록 데이터가 직접 뒷받침하는 내용과 그 범위를 넘어 추론한 내용을 더욱 신중하게 구분해야 합니다.
데이터는 도구이며, 모든 도구는 세계의 일부 측면을 더 쉽게 인식하게 하는 동시에 다른 측면을 덜 보이게 합니다. 측정은 인간의 지식을 확장하지만, 특정한 방법이 드러낼 수 있는 대상의 경계도 설정합니다.
현대적 적용: 정보 과잉 시대의 데이터 리터러시
현대사회는 수치·텍스트·시각·행동 데이터가 끊임없이 흘러드는 환경에 놓여 있습니다. 정부와 기업, 과학자와 인공지능 시스템은 이러한 자료를 사용하여 패턴을 찾고, 대안을 비교하며, 예측합니다. 데이터 기반 분석은 의사결정을 개선할 수 있지만, 그 결과는 여전히 기반 데이터가 생산된 조건에 의존합니다.
그렇다면 우리는 무엇을 해야 합니까? 첫째, 데이터의 출처뿐 아니라 방법론도 검토해야 합니다. 둘째, 데이터가 포착하지 못한 모집단·변수·경험·불확실성을 찾아야 합니다. 셋째, 데이터에 기반한 결론을 절대적 진리가 아니라 규정된 한계 안에서 뒷받침되는 주장으로 다루어야 합니다.
이러한 태도는 모든 숫자를 의심하는 것이 아닙니다. 책임 있는 해석과 수치적 권위에 대한 수동적 복종을 구분하는 것입니다. 데이터를 현실로 취급하는 사람은 자신이 검토한 적 없는 범주의 지배를 받을 수 있습니다. 데이터를 측정된 표상으로 이해하는 사람은 그 한계를 평가하고, 다른 형태의 증거와 비교하며, 데이터에 어느 정도의 권위를 부여할 것인지 결정할 수 있습니다.
데이터는 현실을 완전하게 비추는 거울이 아닙니다. 관찰과 측정, 분류와 해석을 통해 구성된 렌즈입니다. 어떤 렌즈는 다른 렌즈보다 더 정확하고 신뢰할 수 있지만, 어떤 렌즈도 세계의 모든 측면을 한 번에 보여주지는 않습니다.
렌즈의 곡률을 이해할 때에만 데이터가 부분적으로 표상하는 현실에 대해 더욱 정당한 해석을 형성할 수 있습니다. 데이터가 어떻게 생산되는지를 계속 질문하는 한, 데이터는 사고를 대체하는 권위가 아니라 사고를 지원하는 도구로 남을 수 있습니다.
참고문헌:
후설, E. (1982). '순수현상학과 현상학적 철학의 이념들' 제1권(Ideas Pertaining to a Pure Phenomenology and to a Phenomenological Philosophy, First Book). F. 케르스텐 번역. 마르티누스 네이호프 출판사.
아이디, D. (1990). '기술과 생활세계: 정원에서 지구로'(Technology and the Lifeworld: From Garden to Earth). 인디애나대학교 출판부.
하이젠베르크, W. (1958). '물리학과 철학: 현대과학의 혁명'(Physics and Philosophy: The Revolution in Modern Science). 하퍼 앤드 브라더스.
데밍, W. E. (1994). '산업·정부·교육을 위한 새로운 경제학'(The New Economics for Industry, Government, Education). 제2판. MIT 출판부.
포터, T. M. (1995). '숫자에 대한 신뢰: 과학과 공적 생활에서 객관성의 추구'(Trust in Numbers: The Pursuit of Objectivity in Science and Public Life). 프린스턴대학교 출판부.
보커, G. C., 스타, S. L. (1999). '분류하기: 분류 체계와 그 결과'(Sorting Things Out: Classification and Its Consequences). MIT 출판부.
이 글은 에드문트 후설의 지향성 분석, 돈 아이디의 기술적 매개에 관한 철학, 베르너 하이젠베르크의 양자 측정 한계에 관한 설명, W. 에드워즈 데밍의 조작적 정의와 측정 체계 이론, 시어도어 포터의 정량적 객관성에 관한 역사 연구, 제프리 보커와 수전 리 스타의 분류 체계 분석에 기반합니다. 대규모 언어 모델·데이터 리터러시·현대의 데이터 기반 의사결정에 관한 논의는 이러한 이론적 틀을 확장한 것이며, 해당 문헌에서 직접 도출한 내용은 아닙니다.
Previous Episode - The Influence of Quantum Mechanics and the Uncertainty Principle on Modern Philosophy
Related Label - Science
Next Episode - Why Is a Model Not Reality? Scientific Models and the Limits of Understanding the World

Comments
Post a Comment