From Manifestation to Latent Status An Anti-Goodhart Architecture for AI Evaluation and Human Consciousness
Zenodo (CERN European Organization for Nuclear Research) July 15, 2026 DOI: 10.5281/zenodo.21381207 via OpenAlex
Summary
AI-generated from the abstractObservable behaviors in artificial intelligence and human consciousness research are often mistakenly taken as proof of deeper internal properties, such as interpreting fluent self-description as genuine subjectivity. This paper introduces a claim–evidence discipline requiring domain-specific evidence to justify moving from observable manifestations to claims about latent states, using two ordinal maps to identify unjustified leaps. The framework integrates preregistration, sham conditions, negative controls, and other safeguards, and is applied separately to AI evaluation and human-consciousness research without treating them as equivalent. A worked feasibility protocol and construct-existence test illustrate how to empirically study lived presence. The main contribution is a transferable methodological architecture for preventing unjustified transitions from what a system displays to what it is claimed to contain.
Study at a glance
| Characteristics | Theoretical or philosophical paper Peer reviewed |
|---|---|
| Keywords | Cognitive architecture Process computing Cognition Terminology Consciousness |
| Key finding | Observable manifestations are often treated as sufficient evidence of deeper latent properties, and a claim–evidence discipline with domain-specific evidential bridges is needed to prevent unjustified transitions from what a system displays to what it is claimed to contain. |
Abstract
This methodological perspective addresses a recurring inferential problem in the study of artificial intelligence and human consciousness: observable manifestations are often treated as sufficient evidence of deeper latent properties. Fluent self-description may be interpreted as subjectivity, repeated behavior as an internal commitment, stable output as intrinsic direction, cognitive access as phenomenal presence, and agreement between observers as objectivity. The paper consolidates a multi-year research program developed across temporal organization, Anti-Goodhart evaluation, black-box directional testing, internal directional development, relational objectivity, and projected presence. Rather than proposing a shared ontology for biological and artificial systems, it introduces a claim–evidence discipline: every transition from an observable manifestation to a stronger latent-status claim requires a domain-specific evidential bridge capable of distinguishing that claim from plausible lower-level explanations. Two conceptual ordinal maps—claim depth and evidence depth—are used to identify unjustified “diagonal jumps” from shallow observations to deep conclusions. The framework integrates preregistration, sham conditions, negative controls, terminology stripping, cross-context transfer, hidden holdouts, intervention-sensitive testing, explicit claim ceilings, and mandatory downgrade verdicts. It also addresses adaptive systems that may learn, anticipate, or game the evaluation process itself. The paper applies this architecture to two distinct domains without treating them as ontologically equivalent. In AI evaluation, it separates output behavior, runtime persistence, training-level tendencies, mechanistic representations, causal direction, and consciousness. In human-consciousness research, it separates organismic regulation, cognitive access, report, metacognition, self-model, and lived presence. A worked Stage IA feasibility protocol and a prospective Stage IB construct-existence test illustrate how the framework can be translated into an empirical program for lived presence. The strongest current contribution is methodological: a transferable architecture for preventing unjustified transitions from what a system displays to what it is claimed to contain.