In-depth practical guides

Keep a reproducible version of public data

Connect source file, dates, schema and transformations: retain data that can be located, compared and reused without confusing an update, correction or change of scope.

Updated :

Shelves in Ifremer photographic archives, photographed by Stephane Lesbats.

Shelves in Ifremer photographic archives, photographed by Stephane Lesbats. Documentary context photograph; it does not depict an Alexis Roux ZG assignment or equipment. © Stephane Lesbats / Ifremer · Source · CC BY 4.0. Resized and converted to WebP; cropped display.

Can another reader reconstruct the result?

A link to a changing database does not identify the data actually used. Corrections may revise old series, exports may lose leading zeros and new columns may change the meaning of a total. A useful version explains the starting material and processing path. The aim is reproducibility of the described work, not a universal guarantee of source quality.

  1. Define scope

    Name producer, dataset, territory, population, period and filters. Separate observation, publication and download dates. An export received today may describe an old period. State the question and exclusions important to interpretation.

  2. Retain the actual input

    Record URL or public parameters, retrieval date and received file when permitted. Assign a version identifier and compute a file digest. A digest helps distinguish files; it does not establish value accuracy. Keep license information with the version.

  3. Describe the schema

    Document fields, types, units, keys and missing values. For CSV, inspect delimiter, encoding, quotes and line breaks in the actual export. Preserve identifiers requiring leading zeros as text. Distinguish zero, missing and not-applicable values.

  4. Trace transformations

    Record filtering, normalization, joins, conversions and deduplication rules in order. Separate raw input from prepared output. Retain join keys and unresolved cases. Similar names alone do not justify merging entities. Document choices instead of hiding them in the final file.

  5. Compare versions

    Measure added, removed and changed rows using a stable key where available. Also inspect structure, units and definitions. Row-count changes do not establish changes in the underlying phenomenon: filters, coverage or producer revisions may explain them.

  6. Prepare reuse

    Associate the result with its source version, steps and limits. Check links leading back to the producer and context. Plan a new collection with comparison criteria. Retain an older version separately when it supports a published conclusion, respecting conditions and the data it contains.

Situations and decisions

Typical situations for preparing a check. They do not describe completed assignments or actual observations.

Identifiers change after opening

Software treats a column as numeric and removes leading zeros. Return to the raw export, require text for the key and check joins. Reformatting an already altered value may not reconstruct the original identifier.

A historical series is revised

The new file also replaces earlier values. Compare by period and key, retain both versions and read producer notes. Separate revised observations from newly added ones before describing a trend.

The file grows without a measured improvement

Coverage now includes more territories or categories. Compare the common scope, then describe the extension separately. Record field changes and missing values rather than assuming every row is comparable.

Record to retain

  • Producer, URL, public parameters, license and separate dates.
  • File version, digest, schema, units and missing-value conventions.
  • Ordered transformations, join keys and unresolved cases.
  • Version comparison, reconstructed result and interpretation limits.

Official references

Related protocols

Checks before reaching a conclusion

  • Does an identifier preserve its exact characters after import and export, including leading zeros?
  • Does the comparison separate changes in values, schema and coverage instead of reducing every difference to a row count?
  • Can another reader reconstruct the output from the retained version and the documented transformation order?

Frequently asked questions

Should every export be retained?

Keep versions needed for the work and published conclusions under a defined retention rule. Avoid purposeless copies and respect storage and reuse conditions.

Does a digest guarantee quality?

No. It identifies file content and helps detect a change; it does not validate provenance, interpretation or values.

Can rows be merged by name?

Check keys, context and namesakes first. A matching rule must handle ambiguous cases and record merges.

How should corrected data be displayed?

Identify the version and relevant correction, separating the revised result from the previously published one. Do not present a revision as a new observation.