Provenance
Stated source, collection context, ownership chain and available supporting documentation.
Existing datasets can shorten collection timelines, but only when the data is technically usable, sufficiently documented and permitted for the intended application. AMSYNK evaluates available information before representing a dataset as project-ready.

Stated source, collection context, ownership chain and available supporting documentation.
Permitted uses, transfer rights, restrictions and documentation relevant to the proposed project.
Format, duration, duplication, signal or image quality, metadata coverage and sample consistency.
Language, domain, geography, task, demographic and acceptance requirements compared with the available data.
Acceptance rules vary by project, so the controls below are configured against the approved specification rather than treated as universal thresholds.
Existing data can shorten a project only when its rights, technical properties and actual content match the intended use. A filename list or headline hour count is not enough.
Identify where the data originated, what permissions or licenses apply, whether redistribution is allowed and whether the intended AI use is compatible with those rights.
Inspect real files for modality, channel or view layout, codecs, resolution or sample rate, duration, corruption and other properties relevant to the requirement.
Check whether language, geography, task categories, environments, speaker or participant attributes and metadata coverage actually match the buyer specification.
Quoted volume should not be treated as accepted volume. Duplicate, corrupt, irrelevant or non-compliant files can materially reduce the usable dataset after validation.
If the dataset passes the agreed review, delivery can include the approved data package, available metadata, supporting documentation and an agreed transfer structure. Evaluation does not imply automatic approval.