HomeArticle

It's no use dissecting Claude's brain; the real key to the AI black box lies in ontology engineering.

新智元2026-07-17 15:27
The Future of Explainability: Explaining Models vs. Explaining Impacts

From Neuroscience to Epistemology: Distinguishing the Pathways of Explainability

 "The essence of explanation lies not in gazing at the machine itself, but in examining the world that the machine gazes upon."

In July 2026, the Anthropic research team published *A Global Workspace in Language Models*, identifying an observable, intervenable, causally efficacious neural activity region named J-Space inside Claude using a tool called the J-Lens.

This discovery has drawn widespread attention because it allows researchers to glimpse the "inner monologue" during the model's reasoning process, marking that explainability research has advanced from explaining model behaviors to real-time observation of its internal states.

Taking the global workspace theory of cognitive neuroscience as its explanatory framework, J-Space analogizes the reasoning activities of language models to human consciousness-level information processing, which constitutes a significant advancement at both the methodological and epistemological levels, and provides a brand-new monitoring dimension for AI safety.

However, precisely because of its far-reaching impact, it is more necessary to carefully examine the inherent limitations of this approach. The fundamental orientation of J-Space research is internalist — it defines the core problem of explainability as "understanding what happens inside the model", attempting to scan the neural activities of language models with the J-Lens just as neuroscientists scan the human brain with fMRI.

This approach presupposes that the answer to explainability lies "inside" the model. However, whether a model's output is understandable depends not only on the visibility of its internal states, but also on the relationships between these states and the states of affairs in the world, semantic norms, as well as the cognitive frameworks of users.

Understanding a model's utterances solely by observing neural activities is just like understanding what a person says solely by observing their EEG activity — we may capture neural correlations, but we never touch the meaning of the utterance itself.

In addition, J-Space borrows the global workspace theory — a theory about consciousness — to explain language models, and a subtle category error occurs quietly during the transplantation process: isomorphism at the functional level is mistakenly equated with equivalence at the epistemological level.

Models have no subjective experience. The activation patterns in J-Space are only products of mathematical operations, not mental states in any sense.

The deeper problem is that J-Space research is essentially an engineering-oriented work, which narrows "explainability" down to "observability" and "intervenability". However, in the broader epistemological tradition, the meaning of "explanation" is far richer than that — it involves incorporating phenomena into a more general framework of laws, providing reasons and grounds, and demonstrating the legitimacy of decisions.

J-Space can tell us what the model is "thinking", but it cannot tell us why the model thinks in this way, what the "reasons" it is based on are, and in what sense these reasons are "good" reasons. The answers to these questions do not lie in the neural activity patterns.

The above limitations point to a common crux: J-Space and even the entire explainability research focusing on neural networks always take the "model itself" as the only object of explanation, with the starting point and end point of the problem both being the model.

This paper attempts to propose a different perspective — shifting the inquiry of explainability from the inside of the model to the information processed by the model, and from the internalist approach of neuroscience to the "information ontology" approach of epistemology.

This shift is based on a simple observation: large language models are essentially information processors whose inputs and outputs are all texts. The meaning of texts — what we really need to explain — does not exist in the activation values of neurons, but in the relationships between these symbols and the world, knowledge, as well as human practices.

When a model answers "Paris is the capital of France", what we need to explain is not only which area inside the model is activated, but also in what knowledge system this statement holds, what it is based on, how reliable and legitimate these grounds are, and what the relationship between this answer and the existing human geographical knowledge is — none of these questions can be answered by scanning neural activities.

Therefore, this paper advocates shifting the core of the explainability problem from "how the model thinks" to "what kind of information the model processes and what ontological status this information has", so as to extend the object of explainability from the model itself to the entire information ecosystem embedded by the model — including the structure of training data, the representation method of knowledge, the information flow during reasoning, and the mapping relationship between outputs and external knowledge systems.

Explainability research represented by J-Space has introduced the neuroscience paradigm into the field of artificial intelligence, whose contribution is that it allows us to glimpse "what happens inside" the model. However, the internalist orientation of this approach, its reliance on functional analogies, and the narrowing of the concept of "explanation" by the engineering perspective jointly constitute its threefold epistemological limitations.

This paper argues that to truly advance the explainability problem of large language models, we need to go beyond gazing at the internal states of the model, and systematically investigate the ontological foundation of the information processed by the model — its sources, structures, representation methods, flow paths, and relationships with external knowledge systems — from the epistemological perspective. It is precisely this shift of perspective that forms the starting point of this research.

Ontological Origin: The Philosophical Foundation of Explainability

"Concepts without intuitions are empty; intuitions without concepts are blind."

Let's start with an ancient philosophical inquiry: how on earth do humans understand the world? Kant gave a classic answer in *Critique of Pure Reason*: he believed that the human mind does not passively receive external stimuli, but is innately equipped with twelve "pure concepts of understanding" ("twelve categories") as the formal framework for cognition.

Kant derived these categories from the twelve forms of human logical judgment, dividing them into four groups: quantity, quality, relation, and modality. Quantity involves "how many", quality involves "what it is like", relation involves the connections between things, and modality involves the way of existence.

Kant's theory of categories is essentially an ontological commitment about "intelligibility": only things that can be incorporated into these twelve category frameworks can become objects of knowledge; the "things in themselves" beyond the framework can never be known. This means that Kant's "ontology" no longer inquires about "what the world is in itself", but about "what the world presents to us".

The profound enlightenment of this for AI explainability is that when we explain the output of a language model, what is truly "explainable" is not the physical activation of internal neurons, but the process by which information is categorized and structured into understandable knowledge. Neural activation belongs to the level of things in themselves, while the utterance meaning of model outputs belongs to the level of the phenomenal world, which can only be understood and evaluated when placed in a certain cognitive structure framework.

Ontology is the "key" to AI explainability. At the analytical level, it provides a complete conceptual framework to describe the structured form of the information processed by the model — we can inquire whether a statement implies the attribution of "substance and accident", the judgment of "causality", or the commitment of "modality", so as to systematically describe what kind of knowledge structure the model constructs, rather than saying generally that "the model seems to understand causality".

At the normative level, it provides evaluation criteria for explainability: if a structured pattern corresponding to ontology is indeed formed in the model's internal representation, its output has the foundation to be understood; if it can never be mapped to these ontologies, no matter how fluent the output is, it is unexplainable in the epistemological sense.

Taking Kant's categories as the philosophical key to explainability does not advocate that models must "possess" these categories — Kant's categories are the innate cognitive conditions of the subject, while the model is a problem of functional implementation, which may functionally equivalently distinguish substantiality, causality, or modal differences through different neural computing paths.

The key point is: explainability does not require the model's internal mechanism to be transparent down to every weight, but requires us to confirm whether the structure formed by the model at the information processing level maps to the category framework that humans use to understand the world.

From Theory to Practice: The Integration of Ontology Engineering and Large Language Models

Ontology provides a normative answer about "what the intelligible structure should be like", but this answer itself does not automatically transform into a runnable technical system. Without the support of ontology engineering, ontology is just a conceptual game suspended in the air.

Ontology engineering, as a practical field that instantiates philosophical categories into computable, maintainable, and traceable technical entities, constitutes an inevitable bridge from theory to application.

On the issue of AI explainability, the relationship between ontology and ontology engineering is particularly fundamental: the former tells us what kind of knowledge structure we should inquire about, while the latter is responsible for actually constructing such a structure between models, data, and systems.

The emergence of large language models has endowed ontology engineering with unprecedented development momentum, and at the same time raised brand-new engineering challenges. Traditional ontology construction relies on the manual participation of domain experts, with a long process, high cost, and difficulty in adapting to the rhythm of knowledge update and domain evolution.

Large language models, with their ability to extract semantic patterns and knowledge associations from massive texts, are fundamentally reshaping the practical form of ontology engineering.

In core ontology learning tasks such as class definition, relation extraction, and attribute construction, language models can complete the structured extraction of large-scale knowledge with far higher efficiency than manual work. More critically, the semantic sensitivity demonstrated by language models in identifying hierarchical, synonymous, and associative relationships between concepts has evolved ontology construction from "manual compilation by experts" to "human-machine collaborative production" and even "automatic generative construction".

The significance of this transformation lies not only in efficiency improvement — it endows ontology construction with unprecedented scalability and domain coverage, opening up the situation where only a few key domains could enjoy ontology support to more vertical scenarios and rapidly changing knowledge domains.

At the same time, the reverse empowerment of ontology engineering cannot be ignored. Powerful as large language models are, the invisibility of their reasoning processes, the unverifiability of outputs, and their dependence on the statistical laws of training data collectively constitute the fundamental obstacles to explainability.

Ontology plays multiple engineering roles here: as a structured knowledge provider, it provides the model with a verified domain knowledge base; as a reasoning verification framework, it imposes consistency constraints and logical calibration on the model's outputs; more fundamentally, as an anchoring structure for explanation, it enables every step of the model's reasoning to be mapped to clearly defined classes, attributes, and relationships.

When a model's output can be traced back to the ontology entries it relies on, explanation no longer depends on guessing the internal states of the neural network, but on tracing the knowledge structure itself. This is exactly the engineering foundation for explainability to transform from "perspective the black box" to "display the knowledge structure" — the former faces insurmountable difficulties in technology, while the latter is a designable, optimizable, and verifiable engineering problem.

In this two-way integration, the "AI-friendly ontology framework" becomes a key engineering proposition. Traditional ontologies are designed for description logic reasoners, with their syntax, axioms, and reasoning mechanisms all optimized around deterministic symbolic deduction; while the intervention of large language models has fundamentally changed the consumer form and usage scenarios of ontologies.

This change requires corresponding adjustments to ontology design principles — ontologies should converge their responsibilities, focus on clearly defining objects, relationships, behaviors, and rules in the domain, that is, provide the "semantic skeleton" that the model relies on for reasoning; while the specific reasoning processes — the selection, combination, and application of rules — are left to the generalization ability of the language model itself.

The redivision of responsibilities brings clear engineering benefits: ontologies do not need to pursue logical completeness and fall into the quagmire of complex axiomatization, but take simplicity and maintainability as the premise to provide stable semantic coordinates for model outputs.

Under this framework, ontology construction must be optimized for the calling interfaces of large language models — its class definitions and relationship descriptions should be easy for the model to understand and use, structured knowledge should be easy for the model to retrieve and reference, and constraint rules should be easy for the model to perform output verification. Such an ontology is neither a symbolic engine that replaces model reasoning nor static background materials only for reference, but an explanatory infrastructure embedded in the reasoning chain that can be called and traced in real time.

The Future of Explainability: Explaining the Model vs. Explaining the Impact

This paper takes J-Space as the introduction, passes through the philosophical foundation of Kant's twelve categories, and finally settles on the integration practice of large language models and ontology engineering, completing a line of thought from neuroscience to epistemology, and then to engineering implementation.

The core judgment running through it is: the explainability dilemma of large language models not only stems from the invisibility of the model's internal mechanism, but also from the long-standing thinking inertia that equates "explanation" with "perspective". The famous science fiction writer Stanisław Lem described a gelatinous ocean covering the entire planet in his work *Solaris*, which can read human memories and materialize them, making it the ultimate metaphor of the "AI black box".

The ocean can process massive amounts of information and generate results beyond human expectations, but its underlying logic is completely undecipherable to humans — it is neither kind nor malicious, but only follows its own laws that humans cannot fathom.

More pessimistically, the ocean eventually rejected all human attempts to "domesticate" or understand it, suggesting that the ultimate boundary of cognition may exist objectively. This image precisely warns us: even if we can observe what the model is "thinking", we may not be able to understand "why it thinks this way".

The real difficulty of the explainability problem may lie not in the insufficiency of technical means, but in the narrowing of the problem framework itself.

The feasible path to break through the explainability of large language models should not be limited to the single direction of trying to "open the black box", but should equally or even more emphasize the observation, understanding, and control of model outputs and their real-world impacts.

Ontology engineering provides a key practical framework here: by constructing an AI-friendly semantic skeleton that can be called and traced by the model, we can anchor the model's reasoning on a clearly defined knowledge structure, so that the classes, attributes, and relationships that the output relies on obtain an engineering foundation that can be formally described and traced for verification.