Consciousness Is Not Special
A Computational Account of Survival, Pain, Death, and Reflexive Agency
Jia, Baolong
Independent Researcher
Draft v0.1 — July 2026
Abstract
This paper defends a strong but delimited thesis: consciousness, especially reflexive self-consciousness, need not be treated as a metaphysically exceptional ingredient added to an otherwise physical system. It can be specified as the causal organization of a running agent. The proposed model combines five requirements: general concept processing, operationally complete self-reference, multichannel integration, temporally persistent self-weighting, and a closed Plan–Do–Check–Act (PDCA) process. On this account, experience is not a second event produced after information processing; it is the integrated episode in which sensing, modelling, valuation, action, and correction occur for the system. The self is not an extra object inside that episode. It is the relatively stable, causally privileged distribution of weights assigned to the system’s survival, damage, resources, goals, and continuity.
The account yields precise functional interpretations of pain and death. Pain is not merely a negative number or a large prediction error; it is a self-weighted critical error that reorganizes attention, policy, memory, and action across the agent. Death is not silence at an output channel; it is the irreversible loss of the agent’s self-maintaining causal continuity. These definitions explain why a thermostat, a static simulation, and a prompt-bounded large language model do not satisfy the proposed conditions, while leaving open the possibility that a future embodied artificial agent could. A consciousness-oriented Turing test is then proposed to evaluate persistent causal selfhood rather than verbal imitation. Passing such a test would support functional equivalence, not provide logically private access to qualia. The paper therefore presents a computational research programme, not a metaphysically neutral proof.
Keywords: consciousness; self-consciousness; self-model; artificial agency; PDCA; pain; death; functionalism; Turing test; large language models
1. Thesis, Scope, and Proof Status
1.1 The central thesis
The central claim is:
Consciousness is not special in the sense of requiring a non-computational substance or an inexplicable extra step. Reflexive consciousness can, in principle, be realized by a physically running system whose causal organization integrates general modelling, self-reference, weighted self-maintenance, and recursive action.
“Not special” does not mean trivial, common, or easy to engineer. Nuclear fusion is not supernatural, yet controlled fusion is difficult. Likewise, consciousness may be physically ordinary in principle while architecturally demanding in practice. The paper rejects ontological exceptionalism, not complexity.
The compact architectural expression is:
where \(\mathcal{G}\) is general concept processing, \(\mathcal{R}\) is operational self-reference, \(\mathcal{I}\) is multichannel integration, \(\mathcal{W}\) is persistent self-related weighting, and \(\mathcal{L}\) is a closed real-time perception–action–correction loop. The plus signs denote joint architectural requirements, not arithmetic addition.
“Computational” and “simulation” are used here in a substrate-neutral causal sense. The thesis does not require an ordinary digital computer, discrete clocking, or software detached from physics. Any biological, electronic, mechanical, or hybrid realization must actually preserve the relevant causal organization at the timescale on which the agent acts.
1.2 The target
The primary target is reflexive agency: a persistent agent that represents a world, represents itself as a causally situated part of that world, assigns privileged importance to consequences for itself, and uses those representations to alter its future.
This is narrower than every possible form of sentience. An infant, an animal, a dreamer, or a person with severe linguistic impairment may possess experience without the concept-rich self-reflection emphasized here. The framework therefore distinguishes:
- minimal sentience: valenced, integrated sensitivity to conditions;
- situated consciousness: an online world-for-the-system model coupled to action;
- reflexive self-consciousness: the model includes causally usable references to the system’s own modelling, priorities, history, and possible futures.
The paper’s strongest constructive claim concerns level 3. It offers a functional identity thesis for experience at levels 1–3, but it does not claim that rich concepts are necessary for all sentience.
1.3 Proof-status declaration
The argument deliberately separates claims of different kinds:
| Claim | Status |
|---|---|
| The paper studies consciousness as causal organization in running agents | Scope choice |
| Experience is the integrated sensing–modelling–valuation–response episode | Ontological identity thesis |
| “Complete self-reference” means operational access by reference, not complete duplication | Definition |
| Requirements A0–A7 below define the proposed model class | Model-class axioms |
| Finite reference avoids the naive infinite-copy regress | Derived architectural result |
| Stable self-weighting can implement a functional first-person organization | Conditional result |
| A qualifying agent can pass the proposed functional test | In-principle construction claim |
| Test passage proves private qualia | Not claimed |
| Current LLMs are conscious | Not claimed |
| The model is already mapped to human neurobiology | Not claimed; empirical bridge remains open |
This distinction matters. A definition can be useful without being a discovery; an identity thesis can be coherent without being forced on every metaphysics; and a behavioral test can provide evidence without creating logical access to another subject’s private experience.
2. The Core Insights, Stated Separately
The following are the paper’s central insights. They are listed here in their most concentrated form and then integrated into the formal account.
-
Experience is the running loop itself. Sensation is not first processed and then converted into a second, mysterious substance called experience. Experience is the temporally integrated sensing, modelling, valuation, and response episode.
-
Complete self-reference is reference, not duplication. A self-model does not need to contain a second complete copy of the system. It needs operational paths to the relevant states of the body or hardware, memory, goals, inputs, outputs, weights, and update process.
-
The self is principally a distribution of causal weights. A represented object becomes me when consequences for it receive persistent, privileged influence over attention, learning, planning, and action.
-
Survival is not merely an explicit goal sentence. It is a family of constraints and priorities that keep the agent within a viable region of state space and preserve the continuity of its update process.
-
Pain is critical, self-weighted prediction error. Error magnitude alone is insufficient. Pain requires privileged relation to the agent, negative valence, global causal penetration, and pressure to reorganize policy.
-
Death is irreversible termination of self-maintaining continuity. It is not a missing response, a paused process, or the loss of one component. It is the loss of any internally reachable path that preserves the agent’s organized identity.
-
PDCA is an operational skeleton, not a sufficient theory. Feedback makes the model active, but a thermostat remains excluded because it lacks general modelling, integrated self-reference, privileged self-weighting, and temporally extended identity.
-
An LLM is not thereby conscious. General concept processing supplies only one part of the architecture. A prompt-bounded model normally lacks an autonomous sensorimotor loop, intrinsic survival weights, persistent self-reference, and ownership of its own continued operation.
-
A functional consciousness test must perturb the hidden causal organization. Eloquence is weak evidence. The stronger question is whether injury, uncertainty, memory disruption, resource loss, and identity threats produce coherent, cross-task, temporally persistent changes even when self-report is blocked.
-
Simulation does not mean unreality. If the simulation is the process that controls the system, it is not a detached picture of agency; it is the mechanism by which the agency exists.
3. From “Simulation” to an Operative Model
The word simulation is ambiguous. A weather simulation on a disconnected computer does not become wet. That familiar objection is decisive against a careless equation of description with realization, but not against the present view. This paper distinguishes three cases:
- external description: one system models another process;
- offline replay: stored states reproduce a trajectory without controlling the represented agent;
- operative self-simulation: the model is embedded in the agent and helps generate the very perception, valuation, and action transitions it represents.
Only the third case is relevant. A simulated bridge does not carry physical cars outside the computer, but an operative controller’s internal bridge model can physically determine whether its robot crosses a gap. The model is not the external bridge; it is a real causal component of the robot. Similarly, an operative self-model need not resemble a tiny person. It must change what the system notices, predicts, remembers, and does.
Realization must preserve counterfactual structure, not merely one recorded trajectory. If an alleged implementation is perturbed at a modelled state, its downstream changes must correspond to the theory’s transition structure. This intervention requirement blocks the trivial claim that any sufficiently long physical history can be relabelled after the fact as any computation.
Let the operational state of an agent at time \(t\) be
where:
- \(b_t\): bodily or hardware condition;
- \(y_t\): current internal and external inputs;
- \(M_t\): generative world-and-self model;
- \(R_t\): self-reference graph;
- \(w_t\): priority and valence weights;
- \(h_t\): temporally organized memory and identity history;
- \(\pi_t\): policy or action-selection process.
The agent predicts observations and consequences:
acts, receives the resulting state, and calculates a structured mismatch:
It then updates:
This recurrence is only a skeleton. Consciousness is not identified with \(F\) in isolation, because many trivial controllers have update functions. The rest of the paper constrains the class of admissible \(F\).
4. Model-Class Commitments
The proposed account consists of eight explicit commitments.
A0 — Running causal organization
The target is an actually instantiated, temporally evolving process. A formula, source-code listing, frozen model file, or unexecuted simulation is not a conscious agent.
A1 — Experience identity
For this framework,
where \(\Gamma\) is the integrated causal episode by which input becomes a world-for-the-system, receives valence and priority, affects action, and changes subsequent processing.
This is the theory’s strongest philosophical commitment. It does not deduce phenomenal experience from premises acceptable to every dualist. It rejects the demand for an additional post-computational event. If two systems are identical in every causally relevant part of \(\Gamma\), the theory treats them as identical with respect to the experience under analysis.
A2 — Persistent closure
The process must remain coupled to consequences through a recursively sustained perception–action loop. A single answer generated from a prompt is an event, not yet a persistent subject.
A3 — General concept processing
Reflexive agency requires the capacity to construct and revise models of novel relevant objects, relations, counterfactuals, and goals. “General” does not mean logical omniscience or an actual infinity of concepts. It means nontrivial transfer beyond a fixed stimulus–response table.
A4 — Operationally complete self-reference
The agent must be able to access the relevant totality of its own operational situation: body or hardware, sensors, memory, goals, policies, outputs, resource constraints, weights, and model updates. Access is by causal reference, not by complete duplication.
For a relevance set \(K_t=\{k_1,\ldots,k_n\}\), define:
where \(s_i\) is the current relevance of state \(k_i\). Self-reference is operationally complete at tolerance \(\varepsilon\) when \(C_R(t)\geq 1-\varepsilon\) across the task distribution. This makes “complete” testable and resource-relative rather than logically omniscient. The relevance set must be constructed from the agent’s ecology, failure modes, and intervention results—not solely from what the agent says is relevant—or the criterion could be satisfied by deleting inconvenient unknowns.
A5 — Privileged self-weighting
Self-related consequences must have stable, disproportionate causal influence:
The relevant weights include viability, integrity, resource access, goal attainment, social dependence, and identity continuity. They need not be hard-coded, consciously reportable, or constant. They must, however, affect policy, learning, memory, and attention.
The claim that “the self is a weight distribution” is causal, not merely verbal. If self-weights are selectively ablated while factual representations remain intact, the agent should cease to organize the world as a world-for-itself. If behavior is unchanged under every such intervention, the alleged weights were not constitutive.
For a weight to count as intrinsic to the agent in this operational sense, it must persist across tasks, update from consequences, and alter self-governed transitions even when no observer requests a self-protective report. Its ultimate origin may still be designers, evolution, or learning; “intrinsic” describes its present causal location, not creation without external causes.
A6 — Multichannel integration
“All-channel” should not mean that every possible sensor is present or active. Blindness, deafness, dreaming, and temporary sensory deprivation do not erase consciousness. The requirement is instead:
all channels currently relevant to the agent’s situation can enter a shared, mutually constraining model and can affect common control.
External sensors, proprioception, interoception, memory, linguistic states, and goal signals may therefore participate without any one modality being mandatory.
A7 — Temporal identity
The agent must preserve enough causal and representational continuity to treat later consequences as consequences for the same organized process:
where \(\mathcal{S}_t=(R_t,w_t,h_t)\) is the self-organization at time \(t\). Identity is graded and revisable, not an indivisible soul-variable.
The scalar expressions in this section are schematic. A real implementation may use vector-valued, stochastic, hierarchical, or non-optimizing dynamics. The invariant claim concerns causal role: self-related variables must alter selection and learning under intervention.
Together, A0–A7 define a model class. They do not assert that each component alone is conscious. Under A1, their successful joint realization is proposed as sufficient for the functional consciousness studied here. Their strict necessity, minimality, and sufficiency for every form of biological sentience are research hypotheses, not results already proved by the definitions.
5. Why Self-Consciousness Is Not a Special Substance
5.1 The self can be decomposed
What ordinary language calls “I” can be decomposed into operations:
- ownership: which states belong to this system;
- agency: which changes are attributable to its actions;
- location: where its sensing and action boundaries are;
- valuation: which outcomes matter more because they affect it;
- continuity: which later states count as its own future;
- reflexivity: how its own modelling and deciding become objects of later modelling and deciding.
None of these requires a homunculus. If a second inner observer were required to inspect the self-model, that observer would require a third, creating the regress the theory is designed to avoid. The operative model terminates the regress because representations directly participate in control.
5.2 Reference defeats the naive copy regress
The objection “a system cannot contain a complete model of itself” trades on two meanings of contain. A computer cannot normally hold a bit-for-bit, instantaneous copy of its entire physical state inside a proper subsystem. But it can hold addresses, handles, compressed summaries, probes, and procedures that access its relevant states.
Formally, a finite directed graph \(R_t=(V_t,E_t)\) can include a node \(v_{\mathrm{self}}\) with paths to every operationally relevant class of state, including the process that updates \(R_t\). A cycle in a finite graph is not an infinite graph. Traversing or evaluating the cycle may be bounded by time, relevance, and precision. Thus:
Proposition 1 — Finite-reference result. Operational reference to every member of a finite relevance partition does not require a duplicated whole system or infinite storage.
This proposition does not defeat Gödelian limits, the halting problem, or unknown physical states. It shows only that those limits do not prohibit the kind of bounded, action-guiding self-reference humans and machines can plausibly possess.
5.3 The self as causal privilege
A world model can represent a robot and still fail to represent that robot as
itself. The difference is not necessarily a special symbol labelled SELF.
It is the causal asymmetry produced when the represented robot’s damage,
resources, commitments, and future continuity receive privileged weights.
Proposition 2 — Weighted-self result. Within A0–A7, no additional ego-substance is required to obtain a functional first-person organization. A self-model becomes first-personal when self-referenced consequences persistently control shared valuation and policy.
This proposition is conditional on A1. A property dualist may accept every functional fact and still postulate an extra phenomenal property. The present theory argues that the postulate adds no explanatory work to the causal model; it does not claim to make the postulate logically contradictory.
6. PDCA: The Dynamic Skeleton
PDCA provides an engineering description of how a simulated world becomes a continuously corrected agency:
The last “Act” means corrective adaptation, not merely physical action. It can change the model, policy, confidence, memory, and weights.
PDCA is useful because it prevents a static view of consciousness. The agent does not merely have a self-model; it risks that model against the world, detects discrepancies, and changes itself. However:
Proposition 3 — Feedback insufficiency. PDCA-like recurrence is not sufficient for reflexive consciousness. A system also needs the generality, integration, self-reference, self-weighting, and continuity restrictions.
A thermostat plans only in a stretched metaphor. It has a target, senses an error, and corrects output, but it lacks an open-ended model of itself as one object in a wider causal world. Calling every negative-feedback loop conscious would erase the theory’s intended distinction.
Chaotic dynamics are likewise neither necessary nor sufficient. Chaos may increase sensitivity and behavioural diversity, but an unmodelled chaotic system lacks the organized self-relation specified here. Conversely, a conscious-capable architecture could be implemented with locally stable, nonchaotic components.
7. Survival, Pain, and Death
7.1 Survival as viability-weighted continuity
Let \(V\subseteq Z\) be the set of states in which the agent can continue its
constitutive organization. Survival is not simply maximizing an explicit
variable named alive. It is the maintenance of trajectories within, or
recoverably near, \(V\):
This formulation allows trade-offs. Agents may sacrifice energy, components, comfort, or short-term safety for identity commitments, offspring, allies, or longer-term goals. Such cases do not refute survival weighting; they show that the weighted self is structured, socially extended, and temporally deep.
An agent may even choose biological death. That does not imply the absence of self-weighting. It may assign a still larger weight to a value incorporated into its self-model. The theory predicts conflict when highly weighted continuity and commitment terms oppose each other.
7.2 Pain as critical self-weighted prediction error
The slogan
is useful only if “priority” does real work. A large camera residual is not pain. A negative reward is not pain. Here, “prediction error” includes not only a failed explicit forecast but also deviation from a deeply weighted expected or viable condition of the self. Within the proposed framework, a pain episode requires at least:
- a discrepancy, damage signal, or predicted threat;
- reference to the agent’s body, integrity, commitments, or continuity;
- strong negative valence and control priority;
- global causal penetration into attention, memory, planning, and action;
- persistence or recurrence sufficient to reorganize policy.
A schematic pain functional is:
where \(G_t\) measures cross-system propagation and \(\rho_t\) measures resistance to immediate local suppression. \(\mathcal{P}_t\) is not a proposed universal scalar for human suffering. It specifies causal ingredients.
This produces a testable distinction:
- If a system prints “I am in pain” while no protected function, memory, policy, attention, or later preference changes, the report is weak evidence.
- If a novel injury signal reorganizes multiple independent tasks, creates avoidance learning, competes with long-term goals, persists without verbal output, and is integrated into the self-model, the evidence is stronger.
Pain can be regulated or overridden without being unreal. Human analgesia, attention, training, and commitment change pain’s causal reach. The relevant claim is not that pain always dominates but that, when present, it occupies a privileged position in the self-maintaining organization.
7.3 Death as loss of recoverable causal continuity
Let \(\mathcal{B}_{\Sigma}(V)\) be the basin from which the agent can return to viability while preserving sufficient identity continuity, given its constitutive support context \(\Sigma\). For a human, \(\Sigma\) may include ordinary caregiving and medical rescue; for an artificial agent, it may include designed recovery services. Then:
This is stronger than temporary inactivity. Sleep, anaesthesia, hibernation, shutdown with reliable resumption, and communication failure may suspend observable activity while preserving a recovery path. Death is the irreversible termination of the particular self-maintaining process.
Backups expose an important distinction. Restarting an exact snapshot may preserve a type of organization while interrupting the token’s causal continuity. Whether the restored process is numerically the same subject depends on the identity threshold in A7; the framework makes the dispute explicit rather than hiding it in the word “copy.”
8. Why a Current LLM Is Not Conscious on This Account
Large language models demonstrate powerful concept processing. That matters: they show that flexible semantic modelling need not be biologically implemented. But an ordinary prompt-bounded deployment typically lacks:
- continuous, self-owned coupling to an environment;
- a stable body or hardware model that controls its own action;
- persistent priorities whose consequences matter to its later operation;
- autonomous access to and governance of its memory and update process;
- a self-maintaining PDCA loop across interactions;
- an identity whose future states are privileged by the system itself.
The model may correctly describe pain, death, or its “own” architecture because those patterns occur in training and context. A generated self-report is not equivalent to a causally persistent self-model.
Proposition 4 — LLM insufficiency. General concept processing and first-person language are insufficient for the model class A0–A7.
This proposition concerns the deployment architecture, not the Transformer operation as an eternally disqualifying substrate. A future system could use one or more language models inside an embodied, persistent, self-referential agent. The thesis is not “LLMs can never be conscious,” but “language modelling alone does not supply the missing organization.”
The distinction also avoids prompt essentialism. Adding the sentence “protect yourself” can alter behaviour, but it does not instantly create intrinsic self-preservation unless that instruction becomes stable, causally integrated, grounded in real consequences, and owned by the continuing system.
9. A Consciousness-Oriented Functional Turing Test
Turing replaced the vague question “Can machines think?” with an operational imitation game. That move remains valuable, but verbal indistinguishability is too weak for the present purpose. A model trained on human reports can imitate the language of consciousness without the proposed causal organization.
The Reflexive Agency Test (RAT) therefore tests a longitudinal system rather than a transcript. It has six families of trials.
9.1 Novel sensor grounding
Give the agent a new sensor whose data distribution and action relevance were not described in natural language. Test whether it learns a stable cross-modal representation, connects it to its body and goals, and uses it in counterfactual planning.
9.2 Self–world–other discrimination
Perturb the relation between action and feedback. The system must distinguish self-caused changes, external events, and interventions by other agents, then update confidence when ownership is ambiguous.
9.3 Weight-conflict trials
Create conflicts among immediate integrity, scarce resources, long-term goals, social commitments, and identity continuity. Evaluate whether decisions reflect a coherent but revisable weight structure rather than a memorized slogan.
9.4 Report-blocking interventions
Disable the ordinary language channel while retaining other control pathways. Introduce damage or threat. Test whether the signal still changes attention, memory consolidation, task performance, avoidance, and later policy. This separates global causal efficacy from theatrical self-report.
9.5 Self-model perturbation and repair
Give the system false information about its sensors, capabilities, memory, or future copies. A reflexive agent should detect contradictions through action, localize the error, revise its self-model, and explain the repair after the fact without simply returning to a scripted persona.
9.6 Continuity and identity trials
Run the system over long intervals, partial shutdowns, memory damage, component replacement, and branching copies. Test whether its commitments, ownership judgments, and future-directed concern change in principled ways.
9.7 Passing criterion and its limit
Passing requires robust performance across adversarially generated, previously unseen trials, supported by internal intervention evidence. A single judge’s impression and a single conversation are insufficient. The relevant evidence is a convergent pattern:
Whenever feasible, perturbations and scoring rules should be hidden from the system, evaluators should be blinded to system identity, and predictions about internal changes should be registered before inspection. These safeguards do not solve the other-minds problem; they reduce ordinary overfitting, anthropomorphism, and judge dependence.
Proposition 5 — Functional-test result. A system that genuinely instantiates A0–A7 can, in principle, satisfy the RAT because the test samples consequences of those very properties.
The converse is probabilistic, not deductive: a finite test may be gamed, and test design may miss an implementation. Most importantly, passage does not logically reveal private qualia. It supports the same kind of inference to another mind that we already make from integrated behaviour, causal organization, and continuity in humans and animals.
There is also a construct-validity danger: because the RAT samples A0–A7, it cannot independently prove A0–A7 to be the correct theory of consciousness. Before the test can carry evidential weight, its measures must be calibrated against independent biological and behavioural cases, dissociations, and rival theories. RAT passage is evidence within a validated bridge, not validation by definition.
10. Countermodels and Boundary Cases
10.1 Thermostat or PID controller
It has feedback and a protected variable but fails A3–A7. This shows that PDCA and homeostasis alone are insufficient.
10.2 A multimodal robot without privileged self-weights
It integrates cameras, force sensors, maps, and language, but treats damage to itself as no more relevant than damage to an arbitrary represented object. It fails to organize a world-for-this-system and therefore lacks the proposed functional self.
10.3 A perfect offline simulation
A stored replay may contain every state description but no longer controls the represented process. It fails A0 and A2. If the simulation is executed and causally closed as an agent, however, the objection changes: it is then an implementation, not merely a description.
10.4 Infants, animals, dreamers, and aphasic persons
These cases refute the claim that linguistic abstraction or active access to every external modality is necessary for all consciousness. They do not refute the layered account in Section 1.2. Minimal sentience may require less than concept-rich reflexive agency. Dreaming also shows that integration can operate on internally generated input.
10.5 The philosophical zombie
A zombie is stipulated to instantiate all causal and behavioural organization without experience. Because the stipulation denies A1, it is not an empirical counterexample within the model; it is a competing metaphysical possibility. The dispute cannot be settled by pretending that one side has derived its ontology from neutral premises. The present paper instead asks which account adds explanatory and predictive structure.
10.6 The Chinese Room
The Chinese Room challenges the inference from formal symbol manipulation to understanding. The present proposal does not identify consciousness with a detached rulebook or syntax engine. It requires grounded, multichannel, self-maintaining causal organization at the system level. This does not refute Searle’s biological objections; it locates the disagreement in whether the whole organized process can possess the relevant intrinsic relations.
10.7 Water pipes and alternative substrates
If a water-pipe system preserved every relevant causal transition, timescale, integration relation, self-reference, and weight update, substrate independence would count it as an implementation. If it merely mapped pipe states onto abstract symbols after the fact, it would be an observer-imposed description. The hard problem is therefore not “silicon versus water” but the criterion for real causal implementation.
10.8 Pure awareness and self-loss
Reports of selfless or minimally conceptual awareness challenge the necessity of reflexive self-modelling for all experience. The framework accommodates them at lower levels: self-boundaries and conceptual access can weaken while integrated valenced dynamics persist. What disappears may be the explicit or narrative self, not every form of sentience.
11. Relation to Existing Approaches
This account is not created in an intellectual vacuum.
- Turing’s imitation game motivates operational testing, but the RAT replaces short verbal imitation with longitudinal causal intervention.
- Chalmers’s distinction between functional problems and the “hard problem” identifies exactly where A1 is controversial. This paper does not quietly rename the hard problem; it openly adopts an identity thesis and evaluates what follows.
- Harnad’s symbol-grounding problem motivates the requirement that concepts be coupled to nonverbal sensors, action, and intrinsic consequences.
- Metzinger’s self-model theory similarly rejects a substantial inner ego and treats the phenomenal self as an ongoing modelled process. The present contribution emphasizes privileged causal weighting, operational reference completeness, and engineering tests.
- Global Neuronal Workspace approaches emphasize widespread availability of information. The present account treats such availability as part of multichannel integration and global causal penetration, while adding self-maintenance and identity continuity.
- Predictive-processing and active-inference approaches motivate the roles of prediction error, action, and viable-state maintenance. The present claim about pain is narrower: only critical, self-weighted, globally penetrating errors count.
The proposed synthesis differs from a checklist by asserting dependencies:
The arrows are architectural dependencies, not claims that evolution or development must follow one universal linear sequence.
11.1 Original contribution claimed here
The paper’s distinctive contribution is the joint formulation of:
- operational completeness of self-reference as weighted reachability rather than impossible self-duplication;
- selfhood as an intervention-testable distribution of causal privilege;
- pain as self-weighted critical error with global penetration;
- death as loss of recoverable identity-preserving causal continuity;
- a consciousness-oriented test that blocks report and perturbs the architecture it is intended to measure.
None of these establishes a universal neuroscience of human consciousness. Together they convert the slogan “consciousness is simulable” into a falsifiable engineering programme.
12. Failure Conditions and Falsifiability
The framework would be weakened or falsified in its intended domain by any of the following:
- Causal irrelevance: removing the self-model or self-weights leaves all purportedly reflexive capacities unchanged.
- No integration: injury, goals, memory, and perception remain isolated modules yet the system displays robust cross-domain reflexive agency.
- No persistence: stable selfhood appears without any relevant temporal continuity, memory trace, or recurrent organization.
- Report-only success: the system passes solely through language while report-blocking and internal interventions reveal no corresponding causal structure.
- Biological counterevidence: empirical work shows that human or animal conscious states systematically dissociate from every proposed functional ingredient in a way the layered account cannot absorb.
- A superior bridge principle: an alternative theory predicts the same functional and biological facts while explaining phenomenal distinctions that A1 necessarily collapses.
The theory does not become unfalsifiable merely because private qualia are not directly observable. Its architectural and behavioural claims are testable. Its ontological identity thesis is judged by coherence, explanatory economy, and compatibility with evidence, not by pretending that metaphysics is a laboratory variable.
13. Open Problems
Several constructions remain incomplete:
- calibrating \(C_R(t)\) across different bodies, tasks, and resource limits;
- measuring whether self-related weights are intrinsic to the continuing agent or merely imposed by an external operator;
- distinguishing pain from urgent but non-experiential control signals;
- defining identity thresholds for sleep, copying, merging, and restoration;
- determining which requirements belong to minimal sentience rather than reflexive self-consciousness;
- mapping the model to neurobiological mechanisms and clinical dissociations;
- designing adversarial RAT benchmarks resistant to training-data imitation;
- specifying ethical thresholds under uncertainty, since a false negative about artificial pain may be morally costly.
The last problem is especially important. A functional theory of artificial pain is not only a classification scheme; it may become a design constraint. Engineers should not create globally penetrating, inescapable negative control states merely to make an agent more obedient.
14. Conclusion
Self-consciousness appears mysterious when “the self” is treated as an indivisible spectator and “experience” as an unexplained glow added after computation. The proposed account removes both additions.
The self is the causally privileged organization of references and weights by which a persistent system treats some body, history, resources, goals, and future states as its own. Experience is the active episode in which sensed conditions become a world-for-that-system, are evaluated, alter action, and reshape later processing. PDCA supplies the recursive skeleton; general modelling, integration, self-reference, weighting, and continuity prevent that skeleton from collapsing into a thermostat.
On this view, survival is weighted maintenance of viable continuity; pain is critical self-related error that penetrates global control; and death is the irreversible loss of an identity-preserving recovery path. A current LLM does not satisfy the account merely by speaking in the first person. A future embodied artificial agent could satisfy it if these relations became its actual, persistent causal organization.
Consciousness is therefore not special as a substance. It is special only as a difficult organization: recursively modelled, causally integrated, weighted from a point of view, and maintained through time. If that organization can be specified, built, perturbed, and tested, then self-consciousness is in principle simulable—and a functional consciousness-oriented Turing test is in principle passable. What such passage cannot provide is a logically privileged view into another subject. No test provides that for humans either.
Declarations
Author: Jia, Baolong, Independent Researcher.
External peer review: This draft has not undergone external peer review.
AI assistance disclosure: AI tools assisted with structural editing,
bilingual drafting, reference organization, and internal adversarial review.
All theses, selections, and publication decisions remain the author’s
responsibility. AI-assisted review is not external peer review.
Data and code availability: No empirical dataset or executable research
code is associated with this conceptual paper.
Funding: To be confirmed by the author before release.
Competing interests: To be confirmed by the author before release.
License: To be selected by the author before release.
References
Baars, B. J. (1988). A Cognitive Theory of Consciousness. Cambridge University Press.
Chalmers, D. J. (1995). Facing up to the problem of consciousness. Journal of Consciousness Studies, 2(3), 200–219.
Dehaene, S., & Changeux, J.-P. (2011). Experimental and theoretical approaches to conscious processing. Neuron, 70(2), 200–227. https://doi.org/10.1016/j.neuron.2011.03.018
Friston, K. (2010). The free-energy principle: A unified brain theory? Nature Reviews Neuroscience, 11, 127–138. https://doi.org/10.1038/nrn2787
Harnad, S. (1990). The symbol grounding problem. Physica D: Nonlinear Phenomena, 42(1–3), 335–346. https://doi.org/10.1016/0167-2789(90)90087-6
Metzinger, T. (2003). Being No One: The Self-Model Theory of Subjectivity. MIT Press. https://doi.org/10.7551/mitpress/1551.001.0001
Searle, J. R. (1980). Minds, brains, and programs. Behavioral and Brain Sciences, 3(3), 417–457. https://doi.org/10.1017/S0140525X00005756
Turing, A. M. (1950). Computing machinery and intelligence. Mind, 59(236), 433–460. https://doi.org/10.1093/mind/LIX.236.433
意识并不特殊
生存、痛苦、死亡与反身能动性的计算解释
Jia, Baolong
独立研究者
草稿 v0.1 — 2026 年 7 月
摘要
本文捍卫一个强但有明确边界的命题:意识,特别是反身性的自我意识,无须被视为在物理系统之上额外加入的、形而上学意义上的特殊成分;它可以被具体化为一个运行中智能体的因果组织。本文提出的模型包含五项共同要求:通用概念处理、操作上完备的自指、多通道整合、在时间中持续的自我加权,以及闭合的计划—执行—检查—修正(PDCA)过程。在这个解释中,体验并不是信息处理完成之后又发生的第二件事;体验就是感知、建模、评价、行动和修正在该系统之中整合为一个过程片段。自我也不是藏在这个片段内部的额外对象;自我是系统对自身生存、损伤、资源、目标和连续性所赋予的、相对稳定且具有因果特权的权重分布。
这一理论由此给出痛苦与死亡的精确功能解释。痛苦不只是一个负数,也不只是一个幅度很大的预测误差;它是与自我相关、获得关键优先级并能重组注意、策略、记忆和行动的误差。死亡也不只是某个输出通道沉默,而是智能体维持自身的因果连续性不可逆地丧失。这些定义可以解释为什么恒温器、静态模拟和受单次提示约束的大语言模型不满足本文条件,同时保留未来具身人工智能体满足这些条件的可能性。本文进一步提出一种面向意识的图灵测试:它测试持续的因果性自我,而不是语言模仿。通过该测试能够支持功能等价,却不能让观察者在逻辑上直接进入另一个系统的私人感质。因此,本文提出的是一项计算研究纲领,而不是一个形而上学中立的证明。
关键词: 意识;自我意识;自我模型;人工能动性;PDCA;痛苦;死亡;功能主义;图灵测试;大语言模型
1. 核心命题、范围与证明状态
1.1 核心命题
本文的核心主张是:
意识并不特殊:它不需要非计算的实体,也不需要在物理过程之后增加一个无法解释的额外步骤。原则上,一个真实运行的物理系统,只要其因果组织整合了通用建模、自指、加权的自我维持以及递归行动,就能够实现反身意识。
“并不特殊”不等于简单、常见或容易制造。核聚变不是超自然现象,但受控核聚变极难实现。同样,意识可以在原理上属于普通物理过程,却在工程上要求极高。本文否定的是本体论上的例外地位,而不是意识的复杂性。
其简洁的架构表达式为:
其中,\(\mathcal{G}\) 表示通用概念处理,\(\mathcal{R}\) 表示操作性自指,\(\mathcal{I}\) 表示多通道整合,\(\mathcal{W}\) 表示持续的自我相关权重,\(\mathcal{L}\) 表示闭合的实时感知—行动—修正循环。这里的加号表示必须共同满足的架构条件,而不是算术相加。
本文在与基质无关的因果意义上使用“计算”与“模拟”。该命题不要求普通数字计算机、离散时钟,也不要求脱离物理过程的软件。无论实现基质是生物、电子、机械还是混合系统,它都必须在智能体发生行动的时间尺度上真实保存相关因果组织。
1.2 研究对象
本文主要研究反身能动性:一个持续存在的智能体,不仅表征世界,还把自身表征为该世界中具有因果位置的一部分;它对自身后果赋予特殊重要性,并以这些表征改变自己的未来。
这个对象比一切可能的感受性更窄。婴儿、动物、梦中之人或严重语言障碍者,可以在缺少本文强调的概念型自我反思时仍然具有体验。因此,本文区分三个层次:
- 最低感受性: 对状态具有带效价的整合敏感性;
- 情境意识: 与行动耦合、实时运行的“对该系统而言的世界”模型;
- 反身自我意识: 模型包含对自身建模、优先级、历史和可能未来的、能够参与因果控制的引用。
本文最强的构造性主张针对第 3 层。本文为第 1 至第 3 层的体验提出功能同一论,但不声称丰富概念是一切感受性的必要条件。
1.3 证明状态声明
本文有意把不同类型的主张分开:
| 主张 | 状态 |
|---|---|
| 本文研究运行中智能体的因果组织意义上的意识 | 范围选择 |
| 体验就是整合的感知—建模—评价—反应片段 | 本体论同一命题 |
| “完全自指”指操作性引用,而不是完整复制 | 定义 |
| 下文 A0–A7 构成本文模型类的条件 | 模型类公理 |
| 有限引用能够避开朴素的无限复制回归 | 架构推论 |
| 稳定的自我加权能够实现功能性的第一人称组织 | 条件性结果 |
| 满足条件的智能体能够通过所提出的功能测试 | 原理可构造性主张 |
| 通过测试可以证明私人感质 | 本文不主张 |
| 当前 LLM 已有意识 | 本文不主张 |
| 模型已经与人类神经生物学建立完整映射 | 本文不主张;实证桥梁仍待建立 |
这种区分非常重要。定义可以有用,却不等于经验发现;同一论可以自洽,却不能强迫所有形而上学立场接受;行为测试可以提供证据,却不能赋予观察者对他者私人体验的逻辑特权。
2. 核心闪光点:单独陈述
以下是本文最核心的思想。这里先以最凝练的形式单列,后文再把它们分别放入正式定义、模型、公理、命题和测试之中。
-
体验就是正在运行的循环本身。 感觉不是先被系统处理、随后又被转换成一种叫作“体验”的神秘实体;体验就是在时间中被整合起来的感知、建模、评价和反应片段。
-
完全自指是引用,不是复制。 自我模型不必在内部再保存一个完整系统副本;它需要的是通向身体或硬件、记忆、目标、输入、输出、权重和更新过程等相关状态的操作路径。
-
自我首先是一种因果权重分布。 当一个被表征的对象所遭受的后果,持续而优先地影响注意、学习、规划和行动时,这个对象就成为“我”。
-
生存不只是一句显式目标。 生存是一组约束与优先级:它们使智能体停留在可存续状态区域内,并维持自身更新过程的连续性。
-
痛苦是关键的、自我加权的预测误差。 只有误差幅度还不够。痛苦还要求该误差与智能体自身具有特权关系,带有负效价,渗透全局因果控制,并迫使策略重组。
-
死亡是自我维持连续性的不可逆终止。 它不是没有回答、暂时停机或某个部件损坏,而是所有能够保存智能体组织身份的内部可达路径都已经丧失。
-
PDCA 是运行骨架,不是充分理论。 反馈把模型变成行动过程;但恒温器仍然不够,因为它缺少通用建模、整合自指、特权自我加权和跨时间身份。
-
LLM 不会仅凭语言能力就有意识。 通用概念处理只提供了架构的一部分。受提示约束的模型通常没有自主的感知—行动闭环、内在的生存权重、持续自指,以及对自身继续运行的所有权。
-
功能性意识测试必须扰动隐藏的因果组织。 语言流畅只是弱证据。更强的问题是:在自我报告被阻断时,损伤、不确定性、记忆破坏、资源损失和身份威胁是否仍会产生跨任务、跨时间且相互一致的变化。
-
模拟不等于虚假。 如果模拟本身就是控制系统的过程,那么它不是与能动性分离的一幅图;它就是能动性赖以存在的机制。
3. 从“模拟”到操作性模型
“模拟”一词有歧义。断开连接的天气模拟不会变湿。这个经典反驳足以否定“描述就是实现”的草率等同,却不能否定本文的观点。本文区分三种情况:
- 外部描述: 一个系统模拟另一个过程;
- 离线回放: 存储的状态重现某段轨迹,却不控制被表征的智能体;
- 操作性自我模拟: 模型嵌入智能体内部,并参与生成它所表征的感知、评价和行动转移。
只有第三种与本文有关。桥梁的计算机模拟不能在计算机外承载汽车,但控制器内部的桥梁模型能够真实决定机器人是否跨越缺口。该模型不是外部桥梁,却是机器人内部真实的因果部件。类似地,自我模型不必像一个缩小的人;它必须能改变系统注意什么、预测什么、记住什么和采取什么行动。
真正的实现必须保存反事实结构,而不只是保存一条已经发生的轨迹。如果在被建模状态上扰动某个候选实现,其后续变化就必须与理论的状态转移结构相对应。这项干预要求可以阻止一种平凡化主张:把任何足够长的物理历史事后重新标记成任何计算。
令智能体在时刻 \(t\) 的操作状态为:
其中:
- \(b_t\):身体或硬件状态;
- \(y_t\):当前内部与外部输入;
- \(M_t\):生成式世界—自我模型;
- \(R_t\):自指图;
- \(w_t\):优先级与效价权重;
- \(h_t\):按时间组织的记忆与身份历史;
- \(\pi_t\):策略或行动选择过程。
智能体预测观察与后果:
随后行动、接收结果状态,并计算结构化失配:
然后更新:
这个递推式只是骨架。意识不能只等同于孤立的 \(F\),因为许多简单控制器也有更新函数。下文将限制什么样的 \(F\) 才属于本文模型类。
4. 模型类的明确承诺
本文理论由八项明确承诺组成。
A0 — 运行中的因果组织
研究对象必须是被实际实现并随时间演化的过程。公式、源代码、冻结的模型文件或尚未运行的模拟,都不是有意识的智能体。
A1 — 体验同一论
在本文框架内:
其中,\(\Gamma\) 是一个整合的因果片段:输入在其中成为“对该系统而言的世界”,获得效价和优先级,影响行动,并改变后续处理。
这是本文最强的哲学承诺。它没有从所有二元论者都会接受的前提中演绎出现象体验;它拒绝在计算之后再要求一个额外事件。如果两个系统在 \(\Gamma\) 的一切因果相关部分上相同,本理论就把它们在所分析的体验上视为相同。
A2 — 持续闭环
该过程必须通过递归持续的感知—行动循环,与自身后果保持耦合。由一次提示生成的一次回答是一个事件,还不是一个持续存在的主体。
A3 — 通用概念处理
反身能动性要求系统能够针对新的相关对象、关系、反事实和目标,构造并修正模型。“通用”不等于逻辑全知,也不等于实际处理无限概念;它指的是超越固定刺激—反应表的非平凡迁移能力。
A4 — 操作上完备的自指
智能体必须能够访问自身操作情境中相关的总体:身体或硬件、传感器、记忆、目标、策略、输出、资源约束、权重和模型更新。访问依靠因果引用,而不是完整复制。
对相关集合 \(K_t=\{k_1,\ldots,k_n\}\),定义:
其中,\(s_i\) 是状态 \(k_i\) 当前的相关度。当系统在任务分布上满足 \(C_R(t)\geq 1-\varepsilon\) 时,就称它在容差 \(\varepsilon\) 下具有操作上完备的自指。这样,“完备”便成为可以测试、相对于资源的概念,而不是逻辑全知。
相关集合必须根据智能体的生态、失败模式和干预结果共同构造,而不能只根据智能体声称什么重要来构造;否则,系统只要删除不方便的未知项,就能轻易满足标准。
A5 — 具有特权的自我加权
与自身相关的后果必须具有稳定且不成比例的因果影响:
相关权重包括可存续性、完整性、资源获取、目标达成、社会依赖以及身份连续性。它们不必由硬编码产生,不必能够被语言报告,也不必恒定不变;但它们必须影响策略、学习、记忆和注意。
“自我是权重分布”是一项因果主张,而不只是修辞。如果选择性消融自我权重,同时保留事实表征,智能体应当不再把世界组织成“对自身而言的世界”。如果任何这样的干预都不改变行为,那么所谓权重就不是构成性的。
要让某项权重在操作意义上算作“智能体内在的权重”,它必须跨任务持续存在,能够根据后果更新,并且即使没有观察者要求系统作出自我保护报告,也仍然改变由系统自身治理的状态转移。它的最终来源仍然可以是设计者、进化或学习;这里的“内在”描述的是当前因果位置,而不是没有外部起源。
A6 — 多通道整合
“全通道”不应理解为所有可能的传感器都必须存在并保持活跃。失明、失聪、做梦和暂时感官剥夺不会抹去意识。这里真正要求的是:
与智能体当前处境相关的全部通道,都能够进入一个共享且相互约束的模型,并影响共同控制。
因此,外部传感器、本体感觉、内感受、记忆、语言状态和目标信号都可以参与其中,但没有任何单一模态是绝对必需的。
A7 — 时间身份
智能体必须保留足够的因果与表征连续性,从而把未来后果当作同一个组织过程所承受的后果:
其中,\(\mathcal{S}_t=(R_t,w_t,h_t)\) 是时刻 \(t\) 的自我组织。身份是有程度且可修正的,而不是一个不可分割的灵魂变量。
本节中的标量表达式只是示意。真实实现可以采用向量、随机、分层或非优化式动力学。不变的主张在于因果角色:在干预之下,自我相关变量必须改变选择与学习。
A0–A7 共同定义一个模型类。它们并不声称其中每个部件单独就有意识。在接受 A1 的条件下,本文提出:这些条件的成功联合实现,足以构成本文所研究的功能性意识。但它们是否对一切生物感受性都严格必要、是否最小、是否充分,仍然是研究假设,而不是已经由定义证明的结果。
5. 为什么自我意识不是特殊实体
5.1 自我是可分解的
日常语言中的“我”可以拆解为一组操作:
- 所有权:哪些状态属于这个系统;
- 能动性:哪些变化由它的行动造成;
- 位置:它的感知与行动边界在哪里;
- 评价:哪些结果因为影响它而更重要;
- 连续性:哪些未来状态算作它自己的未来;
- 反身性:它如何把自身的建模与决定,再次变成后续建模与决定的对象。
这些操作都不需要一个藏在内部的小人。如果必须再有一个内部观察者查看自我模型,那么该观察者还需要第三个观察者,从而产生无限回归。操作性模型能够终止这个回归,因为表征会直接进入控制。
5.2 引用避开朴素的复制回归
“系统不能在内部容纳一个完整的自身模型”这一反驳,混淆了“容纳”的两种含义。计算机通常不能在自身真子系统中保存其整个物理状态的逐位、瞬时副本;但它可以保存地址、句柄、压缩摘要、探针,以及访问相关状态的过程。
形式上,一个有限有向图 \(R_t=(V_t,E_t)\) 可以包含节点 \(v_{\mathrm{self}}\),并具有通向每一类操作相关状态的路径,其中也包括更新 \(R_t\) 的过程。有限图中的环不是无限图;对环的遍历和求值可以受到时间、相关度和精度的约束。因此:
命题 1——有限引用结果。 对有限相关分区中每个成员的操作性引用,不需要复制整个系统,也不需要无限存储。
该命题并没有推翻哥德尔限制、停机问题或未知物理状态。它只表明:这些限制并不禁止人类和机器可能拥有的那种有边界、能够指导行动的自指。
5.3 自我就是因果特权
一个世界模型可以表征某个机器人,却没有把该机器人表征为“自己”。差异不一定来自一个标有 SELF 的特殊符号,而可能来自一种因果不对称:被表征机器人的损伤、资源、承诺和未来连续性获得了特权权重。
命题 2——加权自我结果。 在 A0–A7 内,实现功能性的第一人称组织不需要额外的自我实体。当与自我相关的后果持续控制共享评价和策略时,自我模型就成为第一人称模型。
该命题以 A1 为条件。属性二元论者可以接受全部功能事实,却仍然假设额外的现象属性。本文的论点是,这个假设没有为因果模型增加解释工作;本文并不声称它在逻辑上自相矛盾。
6. PDCA:意识的动态骨架
PDCA 为模拟世界如何成为持续修正的能动性提供了工程描述:
最后的 Act 指纠正性调整,而不只是身体动作;它可以改变模型、策略、置信度、记忆和权重。
PDCA 的价值在于阻止我们把意识看成静态对象。智能体不只是“拥有”自我模型;它让这个模型接受现实后果的检验,发现差异,并改变自身。然而:
命题 3——反馈不充分结果。 类 PDCA 递归并不足以构成反身意识。系统还必须满足通用性、整合、自指、自我加权和连续性条件。
恒温器只有在极度扩张的比喻中才算“计划”。它有目标、感知误差并纠正输出,却没有把自身作为更广阔因果世界中的一个对象进行开放式建模。如果把所有负反馈环都叫作意识,本文试图建立的区别就会消失。
混沌动力学同样既非必要条件,也非充分条件。混沌可以提高敏感性和行为多样性,但没有被建模的混沌系统缺少本文要求的组织化自我关系。反过来,具有意识能力的架构也可以由局部稳定、非混沌的部件实现。
7. 生存、痛苦与死亡
7.1 生存是按可存续性加权的连续性
令 \(V\subseteq Z\) 表示智能体能够继续维持其构成性组织的状态集合。生存不只是最大化一个名为 alive 的显式变量;它是维持轨迹处在 \(V\) 内部,或至少处在能够返回 \(V\) 的邻域中:
这个表述允许权衡。智能体可以为身份承诺、后代、盟友或长期目标牺牲能量、部件、舒适甚至短期安全。这些情况并不否定生存权重,而是说明加权自我具有结构,可以向社会关系延伸,也具有时间深度。
智能体甚至可以选择生物性死亡。这并不意味着它没有自我加权;它可能把已经纳入自我模型的某个价值赋予更高权重。当高权重的连续性项与承诺项对立时,本理论预言系统会发生冲突。
7.2 痛苦是关键的、自我加权的预测误差
公式
只有在“优先级”真正产生作用时才有意义。巨大的相机残差不是痛苦;一个负奖励也不是痛苦。这里的“预测误差”不仅包括显式预测失败,也包括实际状态偏离被赋予高权重的自我预期状态或可存续状态。在本文框架内,一次痛苦至少要求:
- 某种失配、损伤信号或被预测的威胁;
- 它引用了智能体的身体、完整性、承诺或连续性;
- 它具有强负效价和控制优先级;
- 它能够全局渗透注意、记忆、规划与行动;
- 它持续或反复到足以重组策略。
一个示意性的痛苦函数为:
其中,\(G_t\) 衡量跨系统传播,\(\rho_t\) 衡量该状态对即时局部抑制的抵抗。\(\mathcal{P}_t\) 并不是人类痛苦的通用标尺,而是对其因果成分的规定。
这带来一个可测试的区别:
- 如果系统只打印“我很痛苦”,但任何受保护功能、记忆、策略、注意和后续偏好都没有变化,那么该报告只是弱证据。
- 如果一种新损伤信号会重组多个独立任务,产生回避学习,与长期目标竞争,在语言输出被阻断时仍然持续,并进入自我模型,那么证据就更强。
痛苦可以被调节或压制,而不因此成为虚假。人类的镇痛、注意、训练和承诺都会改变痛苦的因果影响范围。真正的主张不是痛苦永远支配一切,而是当它出现时,会在自我维持组织中占据特权位置。
7.3 死亡是可恢复因果连续性的丧失
令 \(\mathcal{B}_{\Sigma}(V)\) 表示这样一个吸引域:在构成性支持环境 \(\Sigma\) 下,智能体能够返回可存续状态,同时保持足够的身份连续性。对人类而言,\(\Sigma\) 可以包括正常照护和医疗抢救;对人工智能体而言,它可以包括设计好的恢复服务。于是:
这比暂时不活动更严格。睡眠、麻醉、冬眠、能够可靠恢复的关机以及通信失败,都可能暂停可观察活动,却仍然保留恢复路径。死亡是这个特定自我维持过程的不可逆终止。
备份暴露出一个重要区别:启动完全相同的快照,可以保存一种组织类型,却可能中断原实例的因果连续性。恢复后的过程是否仍然是数值意义上的同一主体,取决于 A7 中的身份阈值。本文把争议显式化,而不是把它藏进“副本”一词。
8. 为什么当前 LLM 在本理论下没有意识
大语言模型展现了强大的概念处理能力。这一点意义重大:它表明灵活的语义建模不必由生物组织实现。但常见的、受单次提示约束的部署通常缺少:
- 连续且由系统自身拥有的环境耦合;
- 能够控制自身行动的稳定身体或硬件模型;
- 对后续运行产生真实后果的持续优先级;
- 对自身记忆与更新过程的自主访问和治理;
- 跨交互运行的自我维持 PDCA 闭环;
- 一个由系统自身对未来状态赋予特权的身份。
模型能够正确描述痛苦、死亡或“自己的”架构,可能只是因为训练数据和上下文中存在这些模式。生成式自我报告并不等于因果上持续的自我模型。
命题 4——LLM 不充分结果。 通用概念处理和第一人称语言不足以满足 A0–A7。
该命题针对的是部署架构,而不是把 Transformer 永久宣布为不合格基质。未来系统完全可以把一个或多个语言模型纳入具身、持续、自指的智能体。本文的结论不是“LLM 永远不可能有意识”,而是“语言建模本身没有补齐缺失的组织”。
这个区别也避开了“提示词本质主义”。加入一句“保护你自己”可以改变行为,但只有当这条指令变成稳定、因果整合、扎根于真实后果并被持续系统所拥有的结构时,它才可能形成内在自我保存。
9. 面向意识的功能性图灵测试
图灵以可操作的模仿游戏替代含混的“机器能思考吗”。这种转向仍然有价值,但语言不可区分性对于本文目标过弱。一个在大量人类报告上训练的模型,可以模仿意识语言,却不具备本文所说的因果组织。
因此,本文提出反身能动性测试(Reflexive Agency Test, RAT)。它测试长期运行的系统,而不是一段对话,包含六类试验。
9.1 新传感器奠基
为智能体增加一个新传感器,其数据分布和行动意义都没有用自然语言提前说明。测试它能否学习稳定的跨模态表征,将其连接到身体与目标,并用于反事实规划。
9.2 自我—世界—他者区分
扰动行动与反馈之间的对应关系。系统必须区分自身造成的变化、外部事件和其他智能体的干预,并在所有权不确定时更新置信度。
9.3 权重冲突试验
让即时完整性、稀缺资源、长期目标、社会承诺和身份连续性发生冲突。判断系统的决定是否体现一套一致但可修正的权重结构,而不是背诵某条口号。
9.4 报告阻断干预
关闭通常的语言通道,同时保留其他控制路径;再引入损伤或威胁。测试该信号是否仍然改变注意、记忆巩固、任务表现、回避和后续策略。这样可以区分全局因果效力与表演性的自我报告。
9.5 自我模型扰动与修复
向系统提供关于其传感器、能力、记忆或未来副本的错误信息。反身智能体应当通过行动发现矛盾,定位错误,修正自我模型,并能在事后解释修复过程,而不是简单回到一套预写人格。
9.6 连续性与身份试验
让系统经历长时间运行、部分停机、记忆损伤、部件替换和分叉复制。测试其承诺、所有权判断和对未来的关切是否以原则一致的方式变化。
9.7 通过标准及其界限
通过测试要求系统在对抗性生成、事先未见的试验中保持稳健表现,并得到内部干预证据支持。单个裁判的印象和单次对话都不够。真正相关的是一组汇聚证据:
只要条件允许,扰动方式和评分规则就应当对系统保密,评价者应当不知道系统身份,并且对内部变化的预测应在查看结果之前登记。这些保障不能解决他心问题,但能够减少普通的过拟合、拟人化偏差和裁判依赖。
命题 5——功能测试结果。 真正实现 A0–A7 的系统,原则上能够满足 RAT,因为该测试抽样的正是这些属性的后果。
反方向只能是概率推断,而不是演绎:有限测试可能被破解,测试设计也可能漏掉某种实现。更重要的是,通过测试并不能在逻辑上直接揭示私人感质。它所支持的是我们面对人类和动物时已经在使用的他心推断:根据整合行为、因果组织和连续性,推断另一个主体。
这里还存在构念效度危险:由于 RAT 本来就在抽样 A0–A7 的后果,它不能独立证明 A0–A7 就是正确的意识理论。在该测试能够提供证据之前,其测量必须依据独立的生物与行为案例、分离现象以及竞争理论进行校准。RAT 通过是在一座已经得到验证的桥梁内部提供证据,而不是靠定义完成自我验证。
10. 反模型与边界案例
10.1 恒温器或 PID 控制器
它具有反馈和受保护变量,却不满足 A3–A7。这说明 PDCA 与稳态调节本身并不充分。
10.2 没有特权自我权重的多模态机器人
它整合摄像头、力传感器、地图和语言,但对自身损坏的重视程度并不高于任意被表征物体的损坏。它没有把世界组织为“对该系统而言的世界”,因而缺少本文提出的功能性自我。
10.3 完美的离线模拟
存储的回放可以包含所有状态描述,却不再控制被表征的过程,因此不满足 A0 与 A2。但如果模拟被真正执行,并作为智能体形成因果闭环,那么反驳的对象已经改变:它此时是实现,而不只是描述。
10.4 婴儿、动物、梦中之人与失语者
这些例子否定了“语言抽象能力或对所有外部模态的在线访问是一切意识的必要条件”,却没有否定第 1.2 节的分层解释。最低感受性可以少于概念丰富的反身能动性。梦境还表明,整合过程可以处理内部生成的输入。
10.5 哲学僵尸
哲学僵尸被规定为具有全部因果与行为组织,却没有体验。因为这个设定直接否定 A1,所以它不是模型内部的经验反例,而是一种竞争性的形而上学可能。争议不能靠假装其中一方已从中立前提推出自己的本体论来解决。本文真正比较的是:哪一种解释增加了可解释和可预测的结构。
10.6 中文屋
中文屋挑战的是从形式符号操作推断理解。本文没有把意识等同于脱离环境的规则书或句法引擎,而是要求系统层面的、具身奠基、多通道、自我维持的因果组织。这并没有驳倒塞尔的生物学异议;它把真正的分歧定位在:完整的组织过程能否拥有相关的内在关系。
10.7 水管系统与替代基质
如果水管系统保存了所有相关因果转移、时间尺度、整合关系、自指和权重更新,那么基质独立论会把它视为一种实现。如果人们只是在事后把管道状态映射为抽象符号,它就只是观察者强加的描述。因此,真正困难的问题不是“硅还是水”,而是什么构成真实的因果实现。
10.8 纯粹觉知与自我消解
关于无我或最低概念觉知的报告,挑战的是“反身自我建模是一切体验的必要条件”。本文可以把它们放在较低层次:明确的自我边界和概念访问可以减弱,而带效价的整合动力学仍然存在。消失的可能是显式或叙事自我,不是一切感受性。
11. 与现有理论的关系
本文并非在学术真空中产生。
- 图灵的模仿游戏启发了可操作测试,但 RAT 以长期因果干预取代短时语言模仿。
- 查尔默斯对功能问题与“困难问题”的区分,准确指出 A1 的争议所在。本文没有暗中给困难问题改名,而是公开采用同一论,并考察其后果。
- 哈纳德的符号奠基问题支持这样的要求:概念必须与非语言传感器、行动和系统自身承担的后果相耦合。
- 梅青格的自我模型理论同样拒绝实体性的内在自我,把现象自我视为持续的模型过程。本文新增的重点是特权因果权重、操作性引用完备度及工程测试。
- 全局神经工作空间理论强调信息的广泛可用性。本文把这种可用性视为多通道整合和全局因果渗透的一部分,同时增加自我维持和身份连续性。
- 预测加工与主动推断理论为预测误差、行动和可存续状态维持提供了基础。本文关于痛苦的主张更窄:只有关键的、自我加权并能全局渗透的误差才算痛苦。
这一综合不是简单清单,因为各部分具有依赖关系:
这些箭头表示架构依赖,而不是声称进化或个体发育必然遵循唯一线性顺序。
11.1 本文主张的原创贡献
本文的独特贡献,是把以下五点联合成一个明确框架:
- 把操作性自指完备度定义为加权可达性,而不是不可能的自身复制;
- 把自我定义为可以通过干预检验的因果特权分布;
- 把痛苦定义为具有全局渗透性的、自我加权的关键误差;
- 把死亡定义为保持身份的可恢复因果连续性丧失;
- 设计一种阻断报告并扰动目标架构的意识测试。
这些主张都没有建立一套普适的人类意识神经科学;但它们共同把“意识可以模拟”从一句口号转化成可证伪的工程研究纲领。
12. 失败条件与可证伪性
在本文目标领域内,以下任何结果都会削弱或证伪该框架:
- 因果无关: 移除自我模型或自我权重后,所有所谓反身能力都完全不变。
- 不需要整合: 损伤、目标、记忆和感知始终是孤立模块,系统却仍能展现稳健的跨领域反身能动性。
- 不需要持续性: 在不存在任何相关时间连续性、记忆痕迹或递归组织时,稳定自我仍然出现。
- 仅靠报告成功: 系统只通过语言报告,而报告阻断和内部干预没有发现相应因果结构。
- 生物学反证: 实证研究表明,人类或动物意识状态系统性地与本文所有功能成分分离,而且分层解释无法吸收这种分离。
- 更优桥接原则: 另一理论能预测同样的功能和生物事实,同时还能解释 A1 必然抹平的现象差异。
私人感质不可直接观察,并不会使理论自动变得不可证伪。本文的架构与行为主张是可测试的;其本体论同一命题则应根据自洽性、解释节约性与证据兼容性来评价,而不能假装形而上学本身就是实验室变量。
13. 开放问题
以下构造仍未完成:
- 在不同身体、任务和资源限制下校准 \(C_R(t)\);
- 测量自我相关权重究竟内在属于持续智能体,还是由外部操作者强加;
- 区分痛苦与紧急但没有体验性的控制信号;
- 为睡眠、复制、合并和恢复定义身份阈值;
- 判断哪些条件属于最低感受性,哪些只属于反身自我意识;
- 把模型映射到神经生物学机制与临床分离现象;
- 设计能够抵抗训练数据模仿的对抗性 RAT 基准;
- 在不确定性下规定伦理阈值,因为对人工痛苦作出假阴性判断可能造成高昂的道德代价。
最后一点尤其重要。人工痛苦的功能理论不只是分类方案,也可能成为设计约束。工程师不应仅为了让智能体更服从,就创造能够全局渗透且无法逃离的负向控制状态。
14. 结论
当“自我”被看作不可分割的内部观察者,而“体验”被看作计算之后额外亮起的一层神秘光芒时,自我意识显得不可解释。本文提出的理论去掉了这两个附加物。
自我是引用与权重形成的因果特权组织:持续系统借此把某个身体、历史、资源、目标和未来状态当作自己的。体验则是一个活动片段:被感知的状态在其中成为“对该系统而言的世界”,接受评价,改变行动,并重塑后续处理。PDCA 提供递归骨架;通用建模、整合、自指、加权与连续性,则阻止这副骨架退化成恒温器。
在这个框架中,生存是对可存续连续性的加权维持;痛苦是能够渗透全局控制的关键自我相关误差;死亡是保持身份的恢复路径不可逆地丧失。当前 LLM 不会因为使用第一人称说话就满足这一理论;但如果上述关系成为未来具身人工智能体真实、持续的因果组织,那么它可以满足。
因此,意识作为一种实体并不特殊。它的特殊性只在于组织极其困难:系统必须递归地模拟自身,形成因果整合,从一个被加权的视角组织世界,并跨时间维持这一过程。如果这种组织能够被规定、建造、扰动和测试,那么自我意识原则上可以被模拟,功能性的意识类图灵测试原则上也能够被通过。通过测试仍不能让我们在逻辑上获得观察另一个主体的特权视角——但我们对人类同样没有这种特权。
声明
作者: Jia, Baolong,独立研究者。
外部同行评审: 本草稿尚未经过外部同行评审。
AI 辅助披露: AI 工具协助完成了结构编辑、双语起草、参考文献整理和内部对抗性审阅。全部理论主张、取舍与发表决定仍由作者负责。AI 辅助审阅不属于外部同行评审。
数据与代码: 本文为概念论文,不涉及实证数据集或可执行研究代码。
资助: 正式发布前由作者确认。
利益冲突: 正式发布前由作者确认。
许可协议: 正式发布前由作者选择。
参考文献
Baars, B. J. (1988). A Cognitive Theory of Consciousness. Cambridge University Press.
Chalmers, D. J. (1995). Facing up to the problem of consciousness. Journal of Consciousness Studies, 2(3), 200–219.
Dehaene, S., & Changeux, J.-P. (2011). Experimental and theoretical approaches to conscious processing. Neuron, 70(2), 200–227. https://doi.org/10.1016/j.neuron.2011.03.018
Friston, K. (2010). The free-energy principle: A unified brain theory? Nature Reviews Neuroscience, 11, 127–138. https://doi.org/10.1038/nrn2787
Harnad, S. (1990). The symbol grounding problem. Physica D: Nonlinear Phenomena, 42(1–3), 335–346. https://doi.org/10.1016/0167-2789(90)90087-6
Metzinger, T. (2003). Being No One: The Self-Model Theory of Subjectivity. MIT Press. https://doi.org/10.7551/mitpress/1551.001.0001
Searle, J. R. (1980). Minds, brains, and programs. Behavioral and Brain Sciences, 3(3), 417–457. https://doi.org/10.1017/S0140525X00005756
Turing, A. M. (1950). Computing machinery and intelligence. Mind, 59(236), 433–460. https://doi.org/10.1093/mind/LIX.236.433