Skip to main content Skip to secondary navigation

Desirable Characteristics of Data Repositories

Main content start

Learn how the Stanford Digital Repository meets U.S. federal government recommendations for data repositories.

The Stanford Digital Repository (SDR) can help you meet federal funding agency data sharing requirements, including required repository characteristics set forth in the United States National Science and Technology Council's Desirable Characteristics of Data Repositories for Federally Funded Research. Below you will find the characteristics set forth in the NSTC document and descriptions of how the SDR addresses each of these areas of need. We have not addressed the additional characteristics for repositories storing human data, because the SDR does not accept these kinds of high risk data.

Organizational Infrastructure

Free and Easy Access

The repository provides broad, equitable, and maximally open access to datasets and their metadata free of charge in a timely manner after submission, consistent with legal and policy requirements related to maintaining privacy and confidentiality, Tribal and national data sovereignty, and protection of sensitive data.

SDR supports providing free, public access to the metadata and content files associated with all deposits. No log in or registration is required to access content in the SDR. Depositors may choose to make content available immediately upon deposit to anyone in the world. Depositors must agree to Terms of Deposit that outline practices and responsibilities for preserving privacy.

Clear Use Guidance

The repository ensures datasets are accompanied by documentation describing terms of dataset access and use (e.g., reuse licenses and need for approval by a data use committee).

Depositors may assign a license from a predefined list provided to them and may provide custom terms of use. Licenses and custom terms of use are clearly displayed on repository pages. A license is not required but is strongly recommended. 

Risk Management

The repository has documented capabilities for ensuring that administrative, technical, and physical safeguards are employed to comply with applicable confidentiality, risk management, and continuous monitoring requirements for sensitive data.

Researchers are directed not to share high-risk data or any other data prohibited by SDR’s Terms of Deposit, including data that may violate any law or infringe others’ contractual, publicity, privacy or confidentiality rights (including those protected by HIPAA or FERPA). 

Libraries staff have documented policies and procedures to maintain the information security and integrity of SDR systems and content.

Retention Policy

The repository provides documentation on policies for data retention.

Stanford University Libraries’ documented retention policy specifies that all content is preserved in perpetuity; data may only be deaccessioned via a documented procedure and written authorization of the University Librarian.

Long-term Organizational Sustainability

The repository has a plan for long-term management of data, including maintaining integrity, authenticity, and availability of datasets; has contingency plans to ensure data are available and maintained during and after unforeseen events.

The SDR is a core service of Stanford University Libraries, the mission of which includes long-term preservation of information. SDR is managed by professional IT and digital preservation staff, according to industry best practices for stewardship. SDR’s stable technical infrastructure, dedicated staffing, and permanent funding through the University provide a reasonable expectation of its long-term sustainability. 

Content is replicated multiple times and stored in geo-diverse locations on different media types, providing long-term data management and data integrity to help ensure content availability in the case of unforeseen events, and in line with both Stanford University’s and the Libraries’ information security, disaster recovery and business continuity plans.

Digital Object Management

Unique Persistent Identifiers

The repository assigns a dataset a citable, unique persistent identifier (PID or DPI), such as a digital object identifier (DOI), to support data discovery, reporting (e.g., of research progress), and research assessment (e.g., identifying the outputs of Federally funded research). The unique PID points to a persistent location that remains accessible even if the dataset is de-accessioned or no longer available.

The SDR assigns a unique persistent identifier and a persistent URL (PURL) to every item deposited. Depositors may also elect to have a DOI assigned to the work. SDR DOIs are registered through DataCite. If content assigned a DOI is no longer available, the PURL for the content will remain accessible as a "tombstone" page containing at minimum a citation for the content. Stanford University Libraries is committed to maintaining the persistence of PURLs and DOIs. SDR content can be versioned; in such cases the DOI will always resolve to the newest version of the work.

Metadata

The repository ensures datasets are accompanied by metadata to enable discovery, reuse, and citation of datasets, using schema that are appropriate to, and ideally widely used across, the communities that the repository serves.

The SDR serves the entire Stanford campus and its many and varied research interests. Because of this, depositors are required to submit a minimal set of metadata that is appropriate across this wide scope of research fields. Depositors are encouraged to include any additional metadata or documentation that would be necessary or beneficial for end users to understand, replicate, or reproduce their work. DataCite's metadata schema is used for the creation and assignment of DOIs, which further enhances discovery. Content is indexed and discoverable in SearchWorks, the Libraries' catalog and is exposed to web crawlers for better web discovery. Metadata can be exported in JSON, Schema.org, and MODS formats.

Curation and Quality Assurance

The repository provides or facilitates expert curation and quality assurance to improve the accuracy and integrity of datasets and metadata.

Libraries staff are available to provide consultations on data curation and quality assurance upon request.

Broad and Measured Reuse

The repository ensures datasets are accompanied by metadata that describe terms of reuse and provides the ability to measure attribution, citation, and reuse of data (e.g., through assignment of adequate and openly accessible metadata and unique PIDs).

Depositors are encouraged to assign a license to their work and have the option to obtain a DOI and provide a preferred citation. Licenses, DOIs, preferred citations, and metadata are displayed on repository pages. The SDR publishes usage metrics (views and downloads) on repository landing pages. Deposits with DOIs are also tracked by Altmetric; Altmetric data is posted automatically to repository pages.

Common Format

The repository allows datasets and metadata to be accessed, downloaded, or exported from the repository in widely used, preferably non-proprietary, formats consistent with standards used in the disciplines the repository serves.

Depositors are encouraged to provide content in open, common, non-proprietary formats and in compliance with disciplinary standards when possible, and to provide alternative open formats for content in proprietary formats. Content can be accessed, downloaded, or exported by end users in the format in which it was uploaded to the repository. Metadata can be exported in JSON, Schema.org, and MODS formats.

Provenance

The repository has mechanisms in place to record the origin, chain of custody, version control, and any other modifications to submitted datasets and metadata.

The SDR tracks all modifications made to the metadata and content files for deposited works. Previous versions of a work with substantial updates remain accessible and downloadable from the PURL page. Every file in the SDR is audited multiple times a year against previously created checksums to identify unauthorized or unintentional modifications.

Technology

Authentication

The repository supports authentication of data submitters. The repository has technical capabilities that facilitate associating submitter PIDs with those assigned to their deposited digital objects, such as datasets.

Depositors must authenticate using Stanford's single sign-on system, and their institutional identity is associated with their deposited objects. Depositors may also enter ORCID iDs for authors and contributors of the work. ORCID iDs are included in DataCite metadata when DOIs are assigned. Works are automatically pushed to ORCID records for Stanford authors who have authorized the university to make such updates.

Long-term Technical Sustainability

The repository has a plan for long-term management of data, building on a stable technical infrastructure and funding plans.

SDR is a base-funded service of the Stanford University Libraries, operated by digital preservation and IT experts in continuing staff roles. SDR system and software maintenance is conducted by these dedicated staff who perform upgrades, security audits, and storage migrations as necessary. The SDR is built with open source software components that are adopted, developed, and supported by a large and rapidly growing number of cultural heritage institutions.

Security and Integrity

The repository has documented measures in place to meet well established cybersecurity criteria for preventing unauthorized access to, modification of, or release of data, with levels of security that are appropriate to the sensitivity of data (e.g., the NIST Cybersecurity Framework).

SDR operates in line with Stanford University’s information and cybersecurity practices, including Stanford’s minimum security requirements for moderate risk servers and applications. A non-exhaustive list of measures includes centralized logging, regular security scans and intrusion detection software, regular and out-of-cycle patching of vulnerabilities, and the use of layered application security including multiple firewalls and all systems being located in an access-controlled data center.