Skip to content

A comprehensive framework for managing the knowledge produced by research institutions

Research institutions produce daily a large amount of knowledge assets: scientific articles, theses, technical reports, datasets, code, teaching materials, project records, patents, working documents, periodical publications, metadata, researcher profiles and bibliographic collections. However, much of that knowledge often remains scattered across isolated platforms, heterogeneous formats or internal circuits that are difficult to access, retrieve and reuse.

This framework proposes a comprehensive vision for managing scientific knowledge throughout its entire life cycle: from creation and editing, through storage and preservation, to global discovery, intelligent analysis and conversion into activable institutional memory.

It is not only about implementing technological tools, but about building an institutional knowledge infrastructure capable of connecting people, processes, data, publications, policies and technologies.

Purpose

The purpose of this framework is to serve as a reference model so that research institutions can:

  • Understand the complete cycle of scientific knowledge management.
  • Assess their current level of technological and organizational maturity.
  • Identify gaps, redundancies or disconnections between platforms.
  • Define intervention priorities according to resources, context and objectives.
  • Plan progressive improvements, from operational actions to strategic transformations.
  • Articulate scientific production with institutional memory, open science, digital preservation and analytical intelligence.

This approach recognizes that each institution starts from a different situation. Some first need to strengthen their editorial platforms or repositories; others require integrating dispersed systems, improving metadata quality, ensuring long-term preservation or adding advanced layers of analysis, knowledge graphs and artificial intelligence.

Therefore, the framework does not propose a rigid sequence, but an evolutionary and modular model.

Guiding principles

The framework rests on several fundamental principles:

  • Open science and responsible access: Knowledge should be visible, accessible and reusable to the greatest extent possible, while respecting ethical, legal, institutional or security restrictions.

  • Interoperability: Platforms should not function as islands. Information should be able to flow between editorial systems, repositories, libraries, researcher profiles, databases and global networks.

  • Identifiable data and publications: The use of persistent identifiers — such as DOI, ORCID, ROR, handles or others — makes it possible to trace scientific production, reduce ambiguity and facilitate connections between authors, works, datasets, projects and institutions.

  • Metadata quality: The retrieval, preservation and analytics of knowledge depend on consistent, standardized and context-rich metadata.

  • Long-term preservation: Scientific knowledge has patrimonial value. Storing it is not enough; its integrity, authenticity and future readability must be guaranteed.

  • Centrality of researchers and support staff: Technology should reduce friction, not multiply administrative burdens. Adopting the model requires training, governance and clear processes.

  • Technological scalability: The framework accommodates interventions ranging from operational tasks on existing platforms to advanced architectures of artificial intelligence and knowledge graphs.


The five stages of the framework

Although the framework is presented in five stages, its real application is not always sequential. In many institutions, the stages overlap, feed back into one another or are addressed in parallel. For this reason, rather than a chain of stages, the framework should be understood as an ecosystem of institutional capabilities.

Stage 1 — Capture, editing and scientific production

The first stage focuses on the moment when knowledge is created, documented, edited and published. It is the entry point of the knowledge management cycle.

In many institutions, this stage includes scientific journals, institutional publishers, preprint platforms, editorial management systems, collaborative authoring tools, digital lab notebooks, early data repositories and data management plans.

What it covers

This stage encompasses:

  • Editorial management and scientific publishing platforms.
  • Systems for journals, books, proceedings, bulletins, reports or preprints.
  • Collaborative writing and academic editing tools.
  • Research data management plans.
  • Assignment of persistent identifiers.
  • Definition of licenses, copyright and access policies.
  • Peer review, editing, layout and publication workflows.
  • Early registration of descriptive metadata.
  • Linking between authors, affiliations, projects, datasets and publications.

Why it matters

If knowledge is not captured correctly at the source, problems are carried through the entire cycle. Common issues include incomplete metadata, duplicate records, absence of persistent identifiers, lack of version control, ambiguous licenses or disconnection between publications and the data that supports them.

Good management at this stage allows knowledge to be born identifiable, locatable, structured and prepared for later integration.

Key maturity elements

An institution is at an advanced level of this stage when it:

  • Has stable, up-to-date and secure editorial platforms.
  • Uses persistent identifiers for authors, institutions, works and datasets.
  • Incorporates data management plans into research projects.
  • Defines clear access, licensing and preservation policies.
  • Reduces researchers' manual workload through automated workflows.
  • Connects editorial production with repositories, profiles and institutional reporting systems.

Examples of possible actions within this stage

Within this stage there may be very diverse interventions, for example:

  • Updating and maintaining editorial management platforms.
  • Configuring review and publication workflows.
  • Improving metadata in journals or editorial repositories.
  • Implementing DOI, ORCID or other identifiers.
  • Support for editing, standardization and publication quality.
  • Design of preprint policies.
  • Integration between publishing systems and institutional repositories.

This stage frequently constitutes the natural starting point for institutions seeking to organize their scientific production without yet transforming their entire infrastructure.

Stage 2 — Management, storage and institutional memory

The second stage refers to the internal organization of knowledge already produced or in the process of consolidation. Here the institution needs to store, describe, connect and govern its intellectual assets.

This stage involves institutional repositories, digital libraries, research information systems, catalogs, project databases, researcher profiles and scientific data repositories.

What it covers

It includes, among others:

  • Institutional publication repositories.
  • Research data repositories.
  • CRIS or research information systems.
  • Standardized profiles of researchers and groups.
  • Library catalogs and special collections.
  • Descriptive, thematic and administrative metadata.
  • Controlled vocabularies, thesauri and taxonomies.
  • Policies for content intake, validation and curation.
  • Integration between editorial platforms, repositories and evaluation systems.
  • Registration of projects, funding, agreements, theses and derived products.

Why it matters

When an institution does not adequately manage this stage, knowledge becomes fragmented. An article may exist in a journal, a dataset on a local server, a thesis in an isolated repository and the researcher's profile in another incompatible system.

The management and institutional memory stage seeks to transform that dispersion into a connected ecosystem, where each knowledge object can be retrieved, contextualized and related to others.

Key maturity elements

An institution reaches greater maturity at this stage when it:

  • Has inventoried its main knowledge assets.
  • Uses stable and standardized repositories.
  • Applies consistent metadata and controlled vocabularies.
  • Connects publications, datasets, authors, projects and research units.
  • Avoids duplicate or contradictory records.
  • Can generate institutional reports from reliable data.
  • Has defined responsibilities for curation, deposit, validation and maintenance.

Common challenges

Among the most common challenges are:

  • Lack of standardization in author names and affiliations.
  • Incomplete or inconsistent metadata.
  • Repositories disconnected from evaluation systems.
  • Difficulty keeping information up to date.
  • Absence of clear policies on what should be stored, by whom and under what criteria.
  • Excessive burden on libraries, research units or administrative staff.

Examples of possible actions within this stage

Tasks such as the following may be developed at this stage:

  • Audit of repositories and catalogs.
  • Content migration between platforms.
  • Metadata normalization.
  • Integration of repositories with institutional profiles.
  • Implementation or improvement of CRIS systems.
  • Cleaning and deduplication of records.
  • Connection between digital libraries and scientific production.
  • Design of policies for intake and curation of digital objects.

This stage is especially relevant for institutions that already produce and publish knowledge but have not yet managed to turn it into a coherent and usable institutional memory.

Stage 3 — Preservation and long-term digital archiving

The third stage focuses on ensuring that institutional knowledge remains available, intact and usable over time.

Preserving is not simply making backups. Digital preservation involves planning the continuity of digital objects in the face of technological change, format obsolescence, loss of context, storage failures or institutional discontinuity.

What it covers

This stage includes:

  • Digital preservation policies.
  • Long-term digital archiving.
  • Management of sustainable formats.
  • Integrity and authenticity verification.
  • Preservation metadata.
  • Migration and format normalization plans.
  • Backup, redundancy and recovery strategies.
  • Documentation of the technical and administrative context of digital objects.
  • Alignment with reference models such as OAIS and ISO 14721.

Why it matters

Scientific knowledge has not only immediate value, but also historical, patrimonial and legal value. A thesis, a dataset, a technical report or an institutional journal may remain relevant decades after its creation.

If not adequately preserved, the institution risks losing part of its scientific memory due to technological failures, platform changes, system abandonment or lack of documentation.

The OAIS model and responsible preservation

The OAIS model — Open Archival Information System — offers a conceptual basis for designing reliable digital archives. Its logic distinguishes, among others, the following concepts:

  • SIP: information package delivered to the archive.
  • AIP: information package preserved by the archive.
  • DIP: information package distributed to users.

This distinction makes it possible to understand that a received object cannot always be preserved in the same format, structure or level of description in which it was delivered. Preservation requires controlled transformations, documentation and integrity guarantees.

Key maturity elements

An institution advances at this stage when it:

  • Defines which materials must be preserved long term.
  • Establishes clear responsibilities over digital custody.
  • Uses preferred and sustainable formats.
  • Documents technical and preservation metadata.
  • Performs periodic integrity checks.
  • Has disaster recovery plans.
  • Treats preservation as part of the knowledge life cycle and not as an optional later task.

Examples of possible actions within this stage

Interventions such as the following may be carried out at this stage:

  • Diagnosis of the preservation status of repositories and archives.
  • Development of digital preservation policies.
  • Selection of sustainable formats.
  • Definition of backup and custody strategies.
  • Implementation of integrity verification routines.
  • Planning of obsolete format migration.
  • Documentation of long-term archiving procedures.

This stage is usually less visible than publishing or discovery, but it is critical for the patrimonial sustainability of the institution.

Stage 4 — Discovery, access and global networks

The fourth stage is oriented outward: making knowledge visible, retrievable and connectable with broader scientific ecosystems.

It is not enough for the institution to store and preserve its production. That knowledge must be discoverable, citable, linkable and reusable within local, regional and global networks.

What it covers

This stage includes:

  • Institutional search engines and academic metasearch engines.
  • Open access portals.
  • Exposure of metadata through standard protocols.
  • Integration with national and international open science networks.
  • Connectors with aggregators, indexers and databases.
  • Interoperability protocols.
  • Institutional outreach pages.
  • Metrics of use, downloads, visibility and impact.
  • Mechanisms for open access, controlled access or restricted access as appropriate.

Why it matters

The impact of knowledge depends to a large extent on its ability to be found. A publication without adequate exposure, without sufficient metadata or without connection to global networks loses visibility, citations and possibilities of reuse.

This stage turns institutional memory into a public asset, interoperable and connected with the global open science infrastructure.

Key components

Metasearch and discovery portals

They allow simultaneous querying of different repositories, journals, digital libraries and institutional collections.

Interoperability protocols

Standards such as OAI-PMH, REST APIs, RSS, academic sitemaps and open metadata schemas facilitate the harvesting and reuse of information.

Persistent identifiers

The consistent use of DOI, ORCID, ROR and other identifiers improves traceability and avoids ambiguity in authors, institutions and digital objects.

Global knowledge networks

Connection with aggregators, directories and open science networks allows institutional production to become part of broader circuits of scientific circulation.

Key maturity elements

An institution reaches greater maturity at this stage when:

  • Its contents are easily locatable from academic search engines.
  • Its repositories expose interoperable metadata.
  • Its publications are connected with regional or global networks.
  • Its authors have identified and linked profiles.
  • It can measure downloads, visits, alternative impact and reuse.
  • It defines clear access policies according to rights, sensitivities and institutional mandates.

Examples of possible actions within this stage

Tasks such as the following may be developed at this stage:

  • Configuration of harvesting protocols.
  • Improvement of metadata exposure.
  • Integration with indexers and aggregators.
  • Optimization of open access portals.
  • Implementation of federated search engines.
  • Linking with open science networks.
  • Analysis of visibility and impact of digital collections.

This stage is especially valuable for institutions seeking to increase the reach of their scientific production without depending exclusively on commercial platforms.


Stage 5 — Knowledge, intelligence and analysis

The fifth stage represents the most advanced level of the framework: turning institutional memory into an active, analyzable knowledge base oriented to decision-making.

Here the goal is not only to store, preserve or retrieve information, but to activate accumulated knowledge, understand relationships, detect patterns, generate institutional evidence and enable new strategic uses of the scientific heritage.

What it covers

This stage includes:

  • Modeling of institutional knowledge.
  • Knowledge graphs.
  • Linked data and semantic structures.
  • Bibliometric analytics.
  • Institutional dashboards.
  • Analysis of collaboration networks.
  • Detection of thematic strengths.
  • Identification of gaps or emerging areas.
  • Text and data mining.
  • Academic recommendation systems.
  • Private conversational artificial intelligence.
  • Internal assistants over institutional document corpora.

Why it matters

Research institutions accumulate enormous volumes of information, but often fail to turn them into strategic knowledge.

A well-organized institutional memory makes it possible to know what has been produced. A layer of intelligence and analysis goes further: identifying relationships, interpreting trends, anticipating needs, discovering internal capabilities and supporting decisions about research, funding, collaboration, evaluation and impact.

This stage makes it possible to answer questions such as:

  • Which research areas concentrate the greatest institutional production?
  • Which researchers, groups or units collaborate with each other?
  • Which projects have generated the most theses, articles, datasets or patents?
  • Which thematic lines are growing or losing intensity?
  • Which internal knowledge can support a new call, policy or research line?
  • Which gaps exist in institutional production?
  • Which scientific capabilities are distributed across different units but have not been connected?
  • Which information can be queried through a private institutional assistant?

Knowledge graphs

A knowledge graph represents information as a network of entities and relationships. Instead of storing data in isolated tables, semantic connections between objects are modeled.

For example:

  • A publication has as author a person.
  • A person belongs to a research unit.
  • A unit participates in a project.
  • A project produces a dataset.
  • A dataset documents a publication.
  • A publication uses a specific methodology.
  • A methodology relates to a priority institutional line.

This approach enables richer analyses, semantic search, recommendations and visualization of complex relationships.

Private conversational artificial intelligence

An advanced layer of this stage may include conversational assistants contextualized over institutional memory, provided they are implemented with security, privacy, governance and responsibility criteria.

These systems can enable queries such as:

  • "What research has been done on this topic at the institution?"
  • "Summarize the main results of project X."
  • "Which authors have published on this subject?"
  • "What datasets exist related to this research line?"
  • "Which theses address this problem between 2015 and 2025?"

The key is that the AI does not operate over the open internet indiscriminately, but over a controlled, described, versioned and governed institutional corpus.

Key maturity elements

An institution enters an advanced level of this stage when it:

  • Has reliable and standardized metadata.
  • Has integrated its main information sources.
  • Can model relationships between knowledge objects.
  • Uses indicators for strategic analysis.
  • Explores knowledge graphs or semantic systems.
  • Assesses the responsible use of artificial intelligence over its institutional memory.
  • Defines policies of quality, traceability, privacy and audit.

Examples of possible actions within this stage

Interventions such as the following may be developed at this stage:

  • Design of a conceptual model of institutional knowledge.
  • Normalization of entities and relationships.
  • Construction of a knowledge graph.
  • Integration of documentary sources and databases.
  • Advanced bibliometric analytics.
  • Dashboards for decision-making.
  • Prototypes of private conversational AI.
  • Assessment of ethical and governance risks.

This stage is especially relevant for institutions that already have solid foundations and wish to take a leap toward a more strategic, connected and intelligent knowledge management.


Intervention levels

One of the central characteristics of this framework is that it does not require all institutions to implement all stages at the same time. On the contrary, it recognizes different levels of maturity and allows progressive interventions.

This means that an institution can:

  • Start by solving operational problems on existing platforms.
  • Then advance toward the integration of systems and metadata.
  • Later incorporate preservation policies.
  • Subsequently improve its global visibility.
  • And finally explore layers of knowledge, intelligence and analysis.

The framework can be applied both to one-off interventions and to institutional transformation processes of greater scope. In general, actions can be classified into three complementary levels.

1. Operational intervention

Focuses on the day-to-day functioning of platforms and services.

Examples:

  • Updating editorial management systems.
  • Fixing errors in repositories.
  • Content migration.
  • Metadata adjustments.
  • Improvements in publishing workflows.
  • Specialized technical support.

This level is fundamental because it ensures that the existing infrastructure is stable, secure and usable.

2. Tactical intervention

Seeks to organize, standardize and connect components of the ecosystem.

Examples:

  • Metadata normalization.
  • Integration between repositories and institutional profiles.
  • Implementation of persistent identifiers.
  • Definition of intake and curation policies.
  • Connection with open science networks.
  • Information quality audits.

This level allows systems to stop functioning in isolation.

3. Strategic intervention

Aims to transform the institutional capacity to analyze, preserve and activate its knowledge.

Examples:

  • Design of an institutional knowledge graph.
  • Implementation of long-term preservation policies.
  • Institutional memory architecture.
  • Advanced analytics systems.
  • Private conversational AI over document corpora.
  • Dashboards for decision-making.

This level turns knowledge management into a strategic tool for the institution.


Transversal dimensions

For the framework to work, technology alone is not enough. There are transversal dimensions that must accompany all stages.

Governance

There must be clarity about who decides, who publishes, who validates, who preserves and who maintains each component of the ecosystem.

Institutional policies

Guidelines are required on open access, research data, preservation, licenses, ethics, privacy, security and responsible use of artificial intelligence.

Training and adoption

Research, technical, library and administrative staff need to understand the value of the framework and develop competencies to use it properly.

Data quality

Without reliable metadata, there is no robust discovery, effective preservation or trustworthy analytical intelligence.

Sustainability

The model must be financially viable and technically maintainable over time.

Security and privacy

Especially when managing sensitive data, strategic information or artificial intelligence layers over internal documents.


Benefits

An institution that applies this framework can obtain benefits at multiple levels.

For institutional management

  • Greater clarity about what knowledge it produces.
  • Better reporting and evaluation capacity.
  • Reduction of duplicities.
  • Reliable information for decision-making.

For research

  • Greater visibility of scientific production.
  • Better connection between authors, projects and results.
  • More efficient retrieval of prior knowledge.
  • Analytical support for identifying strategic lines.

For libraries and information units

  • Clearer integration of catalogs, repositories and collections.
  • Better organization of institutional memory.
  • Greater preservation and access capacity.

For society

  • Greater access to knowledge.
  • Better traceability of research.
  • More possibilities for scientific and social reuse.
  • Strengthening of open science.

As a diagnostic tool

The framework can also be used as a diagnostic tool. An institution may ask itself:

  • Are its editorial platforms up to date and well configured?
  • Does its scientific production have persistent identifiers?
  • Are its repositories integrated with profiles and internal systems?
  • Is there a clear data management policy?
  • Are the metadata consistent and sufficient?
  • Have digital preservation strategies been defined?
  • Are its contents harvestable by global networks?
  • Can it analyze its production through reliable indicators?
  • Is it prepared to incorporate knowledge graphs or private AI?

The more positive the answers, the greater the institutional maturity in knowledge management.

Knowledge management in research institutions can no longer be limited to maintaining isolated repositories or publishing content without a comprehensive strategy. The current challenge is to build ecosystems capable of capturing, organizing, preserving, connecting, discovering, analyzing and activating scientific knowledge in a sustainable way.

This framework of five stages offers a complete and flexible vision to address that challenge. Its value lies not only in describing technologies, but in offering a roadmap for institutions to evolve from basic operational levels to advanced models of institutional intelligence, respecting their starting point, their capabilities and their strategic objectives.

How to apply the ecosystem in your institution?

Let's talk about your project