worldcat
GUILLERMO AND MICHÈLE DE LA DEHESA LIBRARY CATALOG

What is Research Data?

Research data are the materials, observations, or records collected, generated, or reused during a research project. They provide evidence that supports research findings and allow results to be verified, reproduced, and reused by others.

Research data can include:

  • Laboratory and field notebooks.
  • Survey responses and interview recordings.
  • Images, videos, and audio files.
  • Databases, spreadsheets, and datasets.
  • Software, algorithms, scripts, and code.
  • Models, simulations, and metadata.

The definition of research data may vary across disciplines and projects. Preliminary analyses, draft manuscripts, personal notes, correspondence, and physical specimens are not usually considered final research datasets, but they may still form part of the research record and require appropriate management, retention, or preservation.

1. Types of Research Data

Research data can be classified in several ways:

  • By content: Numerical, descriptive, or visual.
  • By methodology: Qualitative or quantitative.
  • By processing stage: Raw (primary), processed, or analysed.
  • By source: Experimental, observational, or computational.
  • By format: Text (Word, PDF, RTF, etc.), spreadsheets (Excel, CSV, etc.), multimedia (JPEG, MPEG, WAV, etc.), databases (XML, MySQL, etc.), software code (Java, C, etc.), or discipline-specific formats (Mesh, 3D CAD, etc.).
2. Open, Shared, and Restricted Data

Research data can be made available at different levels depending on legal, ethical, and confidentiality considerations.

Open Data

Open Data are research data that anyone can access, use, and share without significant restrictions. To be truly open, data should be accompanied by an open licence that clearly defines how they can be reused. Some licences may require users to provide attribution or to share derivative works under the same terms.

Shared Data

Shared Data are available to specific individuals or groups under defined conditions. Although they may be discoverable, access may be restricted for reasons such as licensing, confidentiality, or non-commercial use. For example, data may be shared only with collaborators or researchers who meet certain requirements.

Restricted Data

Some research data cannot be made openly available because they contain personal, confidential, commercially sensitive, or otherwise protected information. In these cases, access should be controlled through secure environments or data access agreements.

Even when datasets cannot be shared openly, researchers are encouraged to publish metadata so others can discover, cite, and request access to the data where appropriate.

3. Legal Considerations

Research data containing personal information must comply with applicable data protection legislation, including the General Data Protection Regulation (GDPR) and relevant national laws.

In addition, datasets and databases may be protected by copyright or database rights. Researchers should therefore ensure that they have the necessary permissions before sharing or reusing data.

What Is Research Data Management?

Research Data Management (RDM) refers to the responsible organization and handling of research data throughout the entire research lifecycle: from the initial planning of a project to the collection, processing, analysis, storage, sharing, preservation, or secure disposal of its data.

RDM includes the practical, technical, legal, and ethical decisions needed to ensure that research data remain secure, accurate, understandable, accessible, and reusable for as long as necessary.

It applies to all types of research data, regardless of their format, discipline, or whether they can ultimately be made openly available.

1. Why Is Research Data Management Important?

Good Research Data Management helps researchers to:

  • Protect valuable data and reduce the risk of data loss.
  • Organize data consistently and avoid unnecessary duplication.
  • Save time and improve research efficiency.
  • Maintain the quality, integrity, and reliability of data.
  • Protect personal, confidential, or sensitive information.
  • Meet legal, ethical, institutional, and funder requirements.
  • Clarify responsibilities within research teams and collaborations.
  • Improve the transparency and reproducibility of research.
  • Facilitate the discovery, understanding, sharing, and reuse of data.
  • Increase the visibility, impact, and long-term value of research.

RDM also supports the application of the FAIR Principles, which aim to make research data Findable, Accessible, Interoperable, and Reusable. FAIR data are not necessarily open: access may be restricted when required by legal, ethical, contractual, or confidentiality considerations.

2. What Does Research Data Management Include?

Research Data Management covers all the decisions and activities involved in handling data before, during, and after a research project.

Planning

Researchers should identify in advance:

  • What data will be collected, generated, or reused.
  • How the data will be organized, documented, and stored.
  • Which legal, ethical, or security requirements apply.
  • Who will be responsible for managing the data.
  • How the data will be shared, preserved, or disposed of.

These decisions are normally recorded in a Data Management Plan (DMP).

Data Collection and Creation

Data should be collected or generated using consistent methods, appropriate formats, and clearly documented procedures. Researchers should also consider data quality, consent requirements, intellectual property, and any restrictions affecting how the data may be used.

Organization and Documentation

Files should be organized using consistent folder structures, file names, and version-control practices. Data should also be accompanied by sufficient documentation and metadata so that they can be understood and correctly interpreted by the research team and, where appropriate, by other researchers.

Storage, Backup, and Security

Research data should be stored in appropriate and secure systems. Backup procedures, access controls, and security measures should reflect the value and sensitivity of the data. Personal, confidential, or sensitive data may require additional safeguards.

Data Processing and Analysis

Any cleaning, transformation, processing, or analysis performed on the data should be documented. Researchers should preserve data integrity and, where possible, maintain a clear connection between the original data and subsequent versions.

Data Sharing and Access

Researchers should determine which data can be shared, with whom, when, and under what conditions. Data that can be made available should be accompanied by suitable documentation, metadata, access conditions, and a clear licence.

When data cannot be shared openly, their metadata may still be published so that the dataset can be discovered and access can be requested where appropriate.

Preservation and Disposal

Data with long-term value should be prepared for preservation and deposited in a suitable repository or preservation system. This may involve selecting sustainable file formats, creating appropriate metadata, assigning a persistent identifier, and defining access conditions.

Data that should not be retained must be disposed of securely and in accordance with legal, ethical, institutional, and funder requirements.

3. Who Is Responsible for RDM?

Research Data Management is a shared responsibility. All members of a research project should understand their role in creating, organizing, documenting, storing, protecting, and sharing data.

The principal investigator is normally responsible for ensuring that appropriate data management arrangements are in place. However, responsibilities may also be shared among researchers, project partners, data stewards, IT services, research support teams, and the Library.

Roles and responsibilities should be agreed at the beginning of the project and recorded in the Data Management Plan.

Start Planning Your Data Management

Good RDM begins before data are collected. A Data Management Plan (DMP) helps research teams anticipate their data needs, assign responsibilities, identify possible risks, and record how data will be handled during and after the project.

A DMP is not separate from Research Data Management: it is the document that translates RDM principles into practical actions for a particular research project.

Data Management Plans

Good Research Data Management begins with planning. A Data Management Plan (DMP) helps researchers turn RDM principles into practical decisions for a specific research project.

What is a Data Management Plan?

A Data Management Plan (DMP) is a living document that explains how research data will be managed during and after a research project.

It describes the data that will be collected, generated, or reused and sets out how they will be organized, documented, stored, protected, shared, preserved, or securely disposed of. It may also address responsibilities, resources, and any legal, ethical, contractual, or intellectual property considerations affecting the data.

A DMP should be created at the beginning of a project and reviewed regularly. It can be updated as the project develops, particularly when there are changes to the data, research methods, project team, infrastructure, or legal and ethical conditions.

A DMP is not separate from Research Data Management: it records how RDM will be put into practice in a particular research project.

Why is a DMP important?

A Data Management Plan helps researchers to:

  • Identify data management needs and potential risks at an early stage.
  • Establish consistent procedures for organizing, documenting, storing, and protecting data.
  • Clarify roles and responsibilities within the research team.
  • Address legal, ethical, security, and data protection requirements.
  • Anticipate the resources and costs associated with managing data.
  • Plan how data will be shared, licensed, preserved, or securely disposed of.
  • Support research transparency, integrity, and reproducibility.
  • Meet institutional, funder, and ethics review requirements.

DMP requirements vary between funders, institutions, and research programmes. Researchers should therefore check the relevant policies and use the required template, where one is provided.

Even when a DMP is not mandatory, creating one is recommended for any project that collects, generates, or reuses research data.

How to Create a Data Management Plan

Before completing a DMP, researchers should review the requirements of their funder, institution, ethics committee, research programme, and project partners. These requirements may determine the template to use, the information to provide, and when the plan must be submitted or updated.

 

A DMP should be proportionate to the nature, scale, complexity, and sensitivity of the research. The initial version does not need to contain every final detail, but it should identify the main data management decisions and any issues that still need to be resolved.

How to Create a Data Management Plan

A Data Management Plan (DMP) explains how research data will be collected, organized, stored, protected, shared, preserved, or securely disposed of throughout the research lifecycle. The following sections provide a practical guide to the main elements a DMP should address.

1. Describe the Data

Explain:

  • What data will be collected, generated, or reused.
  • Why the data are needed and how they support the research.
  • Their source, type, format, and estimated volume.
  • Whether existing or third-party data will be used.
2. Explain How the Data Will Be Organized and Documented

Describe:

  • File-naming and folder-structure conventions.
  • Version-control procedures.
  • The documentation and metadata that will accompany the data.
  • Any disciplinary standards, vocabularies, or metadata schemas that will be used.
3. Plan Storage, Backup, and Security

Identify:

  • Where the data will be stored during the project.
  • How backups will be created and managed.
  • Who will have access to the data.
  • How data will be securely transferred.
  • Which additional safeguards will be used for personal, confidential, or sensitive data.
4. Address Legal and Ethical Considerations

Explain how the project will manage:

  • Personal or sensitive data.
  • Privacy and confidentiality.
  • Informed consent.
  • Anonymization or pseudonymization, where appropriate.
  • Intellectual property rights and data ownership.
  • Third-party data and permissions.
  • Legal, ethical, contractual, or security restrictions on access and reuse.
5. Define Roles, Responsibilities, and Resources

Specify:

  • Who is responsible for each data management activity.
  • Who will make decisions about access, sharing, preservation, and disposal.
  • Which institutional services or infrastructure will be used.
  • Whether additional staff, storage, software, or funding will be required.
  • How data management costs will be covered.
6. Plan Data Sharing and Access

Describe:

  • Which data will be shared and when.
  • Where the data will be deposited.
  • Whether access will be open, shared, or restricted.
  • Which licences or conditions of reuse will apply.
  • Why any data cannot be shared.
  • How restricted access requests will be managed, where appropriate.

Data should be made as open as possible and as restricted as necessary.

7. Plan Preservation or Secure Disposal

Identify:

  • Which data, documentation, metadata, and software should be preserved.
  • How long they should be retained.
  • Which repository or preservation service will be used.
  • Which file formats are suitable for long-term preservation.
  • Whether a persistent identifier, such as a DOI, will be assigned.
  • Which data should be securely disposed of and how this will be done.

Use a DMP Template

Use a DMP template to record the main decisions about how your research data will be managed throughout the project.

 

You can consult a variety of different DMP templates based on major funders’ requirements via the DMPonline tool, as well as generic templates for personal use. 

 

The template can be adapted to the nature and scale of your research. If your funder, research programme, or institution provides a mandatory DMP template, you should use that version instead.

Review and Update Your DMP

A DMP should be reviewed throughout the research project rather than treated as a one-time administrative requirement.

Update it whenever there are significant changes to:

  • The data being collected, generated, or reused.
  • Research methods or workflows.
  • Project staff, partners, or responsibilities.
  • Storage, security, or technical infrastructure.
  • Legal, ethical, or contractual conditions.
  • Plans for data sharing, repository deposit, preservation, or disposal.

The DMP should also record its version number, date, and a brief description of any major changes.

FAIR Data

FAIR Principles

The FAIR Principles were introduced in 2016 through the publication of the FAIR Guiding Principles for Scientific Data Management and Stewardship. Their goal is to improve the Findability, Accessibility, Interoperability, and Reusability of digital data and other research outputs. A key aspect of the FAIR framework is machine-actionability, meaning that data should be organized in a way that allows computer systems to discover, access, combine, and reuse them with little or no human intervention. This is increasingly important as the amount and complexity of research data continue to grow.

 

The following list shows each principle, together with examples of how to implement them:

FAIR Principles

The FAIR Principles provide guidelines for improving the Findability, Accessibility, Interoperability, and Reusability of research data and metadata.

1. Findable

Data and metadata should be easy for both people and machines to discover after publication through searchable resources and repositories.

  • F1. Assign a globally unique and persistent identifier to both data and metadata.
  • F2. Describe data using rich and detailed metadata.
  • F3. Register or index data and metadata in searchable resources.
  • F4. Metadata should clearly include the identifier of the data they describe.
2. Accessible

Data and metadata should be retrievable through their identifiers using standardized communication protocols.

  • A1. Data and metadata can be accessed through their identifiers using standardized communication protocols.
  • A1.1. These protocols should be open, free, and universally implementable.
  • A1.2. The protocols should support authentication and authorization procedures when necessary.
  • A2. Metadata should remain accessible even if the underlying data are no longer available.
3. Interoperable

Data and metadata should be structured and described using shared standards and community-recognized practices to enable integration, exchange, and reuse across systems.

  • I1. Data and metadata should use formal, accessible, shared, and broadly applicable languages for knowledge representation.
  • I2. Data and metadata should use vocabularies that themselves follow the FAIR Principles.
  • I3. Data and metadata should include qualified references to related data and metadata.
4. Reusable

Data and metadata should be sufficiently described so that they can be reliably reused by others, with clear information about their origin and conditions of use.

  • R1. Data and metadata should be richly described with accurate and relevant attributes.
  • R1.1. Data and metadata should be released with a clear and accessible usage licence.
  • R1.2. Data and metadata should be associated with detailed provenance information.
  • R1.3. Data and metadata should comply with the standards and practices adopted by the relevant research community.
BE CAREFUL!

Data can be FAIR compliant but:

  • There may be errors in the data.
  • Data may be incomplete or may not reflect reality.
  • Data may be obsolete or non-relevant
  • Data may have been obtained against IPR or data protection regulations.
  • Data may not declare what their sources are.

Fairness is a formal indicator, not a quality indicator.

FAIR vs Open Data

FAIR data and Open Data are often confused, but they are not the same concept. Open Data refers to data that can be freely accessed, used, modified, and shared by anyone, typically under an open license that defines the conditions for reuse and attribution.

FAIR Data, on the other hand, follows the principles of being Findable, Accessible, Interoperable, and Reusable. These principles focus on improving data management and stewardship by ensuring that datasets are well-documented, easy to discover, accessible through standardized protocols, compatible with other datasets, and supported by clear licensing and provenance information.

Importantly, FAIR does not mean open. Data can comply with the FAIR principles while remaining restricted due to privacy, ethical, legal, or security concerns. Likewise, a dataset may be openly available but not FAIR if it lacks sufficient metadata, documentation, or standardized formats. In short, Open Data emphasizes free access, whereas FAIR Data emphasizes effective discovery, use, and reuse by both humans and machines.

 

At the same time, FAIR should not be used as a reason to keep data closed when sharing them would be appropriate and necessary. The two ideas also differ in focus: FAIR is mainly about how data are structured, described, and managed so they can be found, understood, and reused, while open data is mainly about legal permission and removing access barriers. Data can be publicly available online and still not be FAIR if they lack metadata, standard formats, or clear provenance. Ideally, data should be both FAIR and open whenever possible, since they address different aspects of responsible data sharing.

Tools

FAIR Data Tools

F-UJI is a web service to programatically assess FAIRness of research data objects at the dataset level based on the FAIRsFAIR Data Object Assessment Metrics.

Link

FAIR-Aware is an online tool which helps researchers and data managers assess how much they know about the requirements for making datasets findable, accessible, interoperable, and reusable (FAIR) before uploading them into a data repository.

Link

Listed here are the seventeen minimum viable metrics proposed by FAIRsFAIR for the systematic assessment of FAIR data objects.

Link

ACME-FAIR helps those managing and delivering relevant professional services to self-assess how they are enabling researchers and their colleagues to do just that, using 7 different guides to explore different issues surrounding challenges to put FAIR principles in practice.

Link

DMP Tools

DMPonline is an online tool that helps researchers create Data Management Plans (DMPs). It provides funder-specific and generic templates, tailored guidance, and sample answers to help meet funder and institutional Research Data Management requirements.

Once completed, your DMP can be downloaded and shared with your research support team for review.

DMPTool is a free, open-source tool that helps researchers create and manage Data Management Plans (DMPs). It provides templates and guidance tailored to the requirements of different funding agencies, as well as the option to create a custom DMP.

With DMPTool, you can:

  • Create, edit, review, and share your DMP.
  • Use funder-specific templates and guidance.
  • Collaborate with other researchers and assign project roles.
  • Add researchers’ ORCID iDs.
  • Request feedback from your institution, where this service is available.
  • Export your DMP in multiple formats, including PDF, Word, text, JSON, CSV, and HTML.

Need Help?

The Library can provide guidance on Research Data Management, Data Management Plans, FAIR data, data documentation, repository selection, licensing, and data sharing.

 

Please contact the Open Access Office at openaccess@ie.edu. We’re here to help!