Whoever governs knowledge also governs Artificial Intelligence.

Artificial Intelligence is rapidly becoming one of the strategic technologies that will shape science, industry, and public institutions over the coming decades. Discussions often focus on which language model performs best or which company will release the next breakthrough. Yet a more fundamental question is emerging:

Who governs the knowledge on which AI systems rely?

Our new paper, published in Earth Science Informatics, argues that trustworthy scientific AI cannot depend solely on increasingly capable language models. It also requires a new way of governing scientific knowledge.

From Language Models to Knowledge Governance

Large Language Models are powerful general-purpose tools. However, when used in scientific contexts, they frequently face a familiar problem: they generate answers based on statistical patterns rather than on institutionally validated knowledge.

Retrieval-Augmented Generation, commonly known as RAG, has become a popular solution because it grounds AI responses in external documents. Nevertheless, simply connecting a model to a collection of documents does not automatically guarantee trustworthy scientific answers.

The quality of an AI system ultimately depends on the quality, provenance, and governance of the knowledge it retrieves.

Researcher-Curated Knowledge Governance

In our work, we introduce the concept of Researcher-Curated Knowledge Governance.

The central idea is straightforward: researchers should not only produce scientific knowledge, but also actively curate, validate, organize, and govern the knowledge that feeds AI systems.

Knowledge therefore becomes an institutional asset rather than an uncontrolled collection of documents.

This approach can make AI systems more transparent, reproducible, and auditable, while preserving scientific quality and institutional responsibility.

The RockGPT Approach

To demonstrate this idea, we developed RockGPT, a platform designed for the geosciences.

RockGPT combines:

  • open-source Large Language Models;
  • Retrieval-Augmented Generation;
  • researcher-curated scientific knowledge;
  • institutional deployment on self-hosted infrastructure.

The objective is not simply to answer questions, but to ensure that answers are grounded in trusted scientific sources selected and maintained by domain experts.

Comparison of average model scores for the baseline and RAG-enhanced configurations

Average model scores for the baseline and RAG-enhanced configurations (Figure 4).

Why This Matters for Europe

The debate on European AI sovereignty often concentrates on developing new foundation models.

Our work suggests that sovereignty extends beyond the model itself.

True sovereignty also means controlling scientific knowledge, infrastructure, governance processes, and data.

This perspective is particularly relevant for European research infrastructures such as EPOS, where FAIR data, trusted scientific resources, and institutional governance already provide strong foundations for building reliable and sovereign AI services.

Rather than merely replicating commercial AI platforms, Europe has an opportunity to build AI ecosystems that reflect its own values: openness, transparency, scientific rigour, institutional accountability, and public responsibility.

The goal may not be to build the largest model.

It may be to build the most trustworthy ecosystem.

Looking Ahead

Researcher-Curated Knowledge Governance is not intended as a replacement for Large Language Models.

Instead, it provides a governance framework that enables these models to become more reliable scientific assistants.

As AI becomes increasingly integrated into research, the decisive question will no longer be only how intelligent a model is, but also who governs the knowledge behind it.

That, we believe, is where trustworthy scientific AI truly begins.

Paper

Researcher-Curated Knowledge Governance for Trustworthy Scientific AI in the Geosciences: the RockGPT Approach

Published in Earth Science Informatics.

https://link.springer.com/article/10.1007/s12145-026-02197-5

Transparency note: the drafting of this blog post was supported by generative artificial intelligence. The content was subsequently reviewed, verified, and approved by the author.