> For the complete documentation index, see [llms.txt](https://docs.libnova.com/labdrive/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.libnova.com/labdrive/data-curation-and-preservation-1/collecting-information-needed-for-re-use-and-preservation.md).

# Collecting Information needed for Re-Use and Preservation

IPELTU uses a very general approach to describing projects, in terms of the what are termed Collection Groups, namely “Initiating”, “Planning”, “Executing” and “Closing” for each requiring Additional Information.

![Phases and cycles in a project which collects/creates information to be preserved/curated](/files/TpeeCkbaBaKFOoge27WE)

The table below provides examples for the various stages. The IPELTU document provides further details and checklists for a number of types of projects.

<table><thead><tr><th width="164">Collection Group</th><th>Initiating</th><th width="150">Planning</th><th width="150">Executing</th><th>Closing</th></tr></thead><tbody><tr><td><p>    </p><p><strong>Additional Information    Area</strong></p></td><td></td><td></td><td></td><td></td></tr><tr><td><strong>Data Object</strong></td><td><ul><li>Estimate of volume of data to be produced</li><li>Ideas of the potential value of the data</li></ul></td><td><ul><li>Update Additional Information from Initiating based on more detailed plans</li><li>Identify types of data (raw, processed, etc.) which should be preserved</li><li>Identify types of data e.g., images, tables – and any generic interfaces</li><li>Quality constraints</li><li>Planned rate of data production</li><li>Expand and add detail</li></ul></td><td><ul><li> Update Additional Information from Planning based on what really happens</li></ul></td><td><p>·* Finalise Additional Information from Executing </p><p>·       Inventory of data produced which should be preserved</p><p>·       Volume that would require preservation</p><p>·       Collect quality checks which may be performed on the data by non-experts</p><p>·       Define Information Properties which may be useful</p><p>·       Checks for (and logs of) any missing data</p></td></tr><tr><td><strong>Representation Information</strong></td><td><p>·       Standards planned to be used</p><p>·       Information Model</p></td><td><p>·       Update Additional Information from Initiating based on more detailed plans</p><p>·       Review applicable standards</p><p>·       Refine Information Model</p><p>·       Choice of data format</p><p>·       Identify Hardware and Software Dependencies</p><p>·       Relationships between data items</p></td><td><p>·       Update Additional Information from Planning based on what really happens</p><p>·       Collect Semantics of the data elements e.g., data dictionaries and other semantics</p><p>·       Collect Format definitions and formal descriptions</p><p>·       Create Other Data Documentation</p><p>·       Calibration and system test tools and system test data that will be delivered</p></td><td><p>·       Finalise Additional Information from Executing </p><p>·       Finalise Representation Information Networks to reasonable level</p><p>·       Identify other software which may be used on the data</p><p>·       Create suggestions for the Designated Community and Representation Information needed</p></td></tr><tr><td><strong>Reference Information</strong></td><td>·       Identify standards which will be used to identify and reference the data and metadata</td><td><p>·       Update Additional Information from Initiating based on more detailed plans</p><p>·       Identify which unique identifiers should be used (e.g., DOI or other)</p></td><td><p>·       Update Additional Information from Planning based on what really happens</p><p>·       Rules, methods, tools for referencing data</p><p>·       Generate references to data as it is being created/captured</p></td><td><p>·       Finalise Additional Information from Executing </p><p>·       Identify what may be used in future to identify the Information</p><p>·       Checks for (and logs of) missing references and logs of any</p></td></tr><tr><td><strong>Provenance Information</strong></td><td>·       Record of origins of the project e.g., in a Current Research Information System (CRI)</td><td><p>·       Update Additional Information from Initiating based on more detailed plans</p><p>·       Define Processing workflow, Processing inputs and Processing parameters</p><p>·       Define System Testing required</p><p>·       Documents from system development milestones</p></td><td><p>·       Update Additional Information from Planning based on what really happens</p><p>·       Documentation about the hardware and software used to create the data, including a history of the changes in these over time</p><p>·       Update Documentation of Processing workflow, Processing inputs and Processing parameters</p><p>·       Record who was responsible for each stage of processing</p><p>·       Record when each stage was performed</p><p>·       Record of any special hardware needed</p><p>·       Record Calibration</p><p>·       Processing logs</p><p>·       Record checking of Fixity</p></td><td><p>·       Finalise Additional Information from Executing </p><p>·       Finalise Provenance handover</p></td></tr><tr><td><strong>Context Information</strong></td><td>·       Outline of background concepts needed to understand the project</td><td>·       Update Additional Information from Initiating based on more detailed plans</td><td><p>·       Update Additional Information from Planning based on what really happens</p><p>·       Collect publications related to the data or the processing system</p><p>·       Potential Value of the data and likely business case for sustainability</p></td><td><p>·       Finalise Additional Information from Executing </p><p>·       Identify related data which may in the future be combined with this data</p></td></tr><tr><td><strong>Fixity Information</strong></td><td> </td><td>·       Fixity mechanism (e.g., CRC or digest) of data which may be preserved</td><td><p>·       Update Additional Information from Planning based on what really happens</p><p>·       Identify any special validation procedures that should be carried out.</p></td><td><p>·       Finalise Additional Information from Executing </p><p>·       Identify how do we verify that all files are intact</p></td></tr><tr><td><strong>Access Rights Information</strong></td><td> </td><td><p>·       What are the restrictions on access in the long term?</p><p>·       Clear identification of Intellectual Property Rights</p><p>·       Owners of the data – who can authorize hand-over</p></td><td>·       Update Additional Information from Planning based on what really happens</td><td><p>·       Finalise Additional Information from Executing </p><p>·       Licenses involved</p><p>·       The owner, and the restrictions on access (licenses), and the intellectual property rights</p></td></tr><tr><td><strong>Packaging Information</strong></td><td> </td><td> </td><td> </td><td><p>·       Details of the way components are packaged together for delivery to a repository</p><p>·       Definition of mechanisms for transferring information to next element in the workflow or next in the chain of preservation (e.g., definitions of SIPs)</p></td></tr><tr><td><strong>Descriptive Information</strong></td><td> </td><td> </td><td><ul><li>Identification of methods for exploration/ quick look at the data</li></ul></td><td><p>·       Finalise Additional Information from Executing </p><p>·       Create browse/query data if needed</p></td></tr><tr><td><strong>Issues Outside the Information Model</strong></td><td><ul><li>Estimated Cost of the project</li></ul></td><td><ul><li>The budget for archiving and its relationship to the.  overall budget for the p</li><li>The schedule for major project milestones and deliveries to the archive.</li><li>Identification of archives which are likely to be able to host the data</li></ul></td><td><ul><li>Update Additional Information from Planning based on what really happens</li></ul></td><td><p>·       Finalise Additional Information from Executing </p><p>·       Schedule of deliveries</p><p>·       Pointers to the components to be transferred to the next element in the workflow or next in the chain of preservation</p><p>·       Potential preservation aims for the information created</p><p>·       Potential risks to preservation and exploitation of the data</p><p>·       Define the mechanism for communication between project and archive.</p><p>·       Define suggested Transformational Information Properties </p><p>·       Publications, or references to publications, including scientific publications, related to the project.</p></td></tr></tbody></table>
