Essential Legal Information
Access key documents, including terms and conditions, privacy policies, and more, to stay informed about Ocean’s legal framework.
Purpose
The purpose of this AI Training Data Policy (the “Policy”) is to establish requirements for the collection, preparation, use, storage, retention and deletion of data used to train, fine-tune, validate, test or otherwise improve AI models developed or operated by REACT.
The objective is to ensure that AI training data is used lawfully, securely, transparently and responsibly, while respecting intellectual property rights, confidentiality, privacy and contractual obligations. REACT applies the principles of purpose limitation, data minimisation, security, traceability, human oversight and responsible AI development.
Scope
This Policy applies to datasets used for AI model training and fine-tuning, machine-learning development, model validation and testing, image recognition and classification, similarity and pattern detection, OCR and visual analysis development, model performance evaluation, and other approved AI development activities.
It applies whether an AI system is developed internally or with an approved technology partner.
Reference Documents
Internal
- POL-AI-01 AI Management Policy
- POL-AI-03 AI Acceptable Use Policy
- PROC-AI-03 AI Human Oversight Procedure
- PROC-AI-04 AI Monitoring Procedure
- PROC-AI-05 AI Incident Management Procedure
External
- ISO/IEC 42001
- ISO/IEC 27001
- EU AI Act
- GDPR
Permitted Training Data
- REACT-owned data where REACT holds sufficient rights for the intended AI use.
- Member-contributed data where the member has expressly authorised use for AI development, training, validation or testing.
- Public-domain data where its status has been reasonably established.
- Openly licensed data where the applicable licence permits the intended AI use.
- Synthetic data, provided its creation does not improperly reproduce protected or confidential source material.
- Publicly accessible data only following appropriate assessment of intellectual property, privacy, contractual, platform and other applicable restrictions.
Prohibited or Restricted Data
The following shall not be included in AI training datasets unless specifically assessed and authorised:
- Confidential member or client information.
- Personal data without an appropriate legal basis and assessment.
- Special-category or highly sensitive personal data.
- Privileged legal communications.
- Credentials, passwords or authentication information.
- Payment or financial account information.
- Material subject to contractual restrictions preventing AI use.
- Copyrighted or otherwise protected third-party material where sufficient rights have not been established.
- Data obtained unlawfully or through unauthorised access.
- Information classified by REACT as unsuitable for AI processing.
Where restricted information is not necessary for the AI purpose, it should be removed, anonymised, masked or otherwise excluded before training.
Member-Contributed Data
Members may voluntarily provide product images, logos, packaging examples, reference materials and other content to improve REACT's AI-supported enforcement capabilities. Appropriate permission covering the intended use must be obtained before such material is used for AI training. Member-contributed data shall be used only for the purposes communicated to and agreed with the member. Appropriate safeguards may include watermarking, access restrictions, dataset separation, controlled storage or other technical measures.
Training Data Approval
Each new dataset must be assessed before being introduced into an AI training environment. The assessment should establish its source and ownership, purpose, applicable rights or permission, presence of personal or confidential information, contractual restrictions, security classification, intended AI system, retention requirements and restrictions imposed by the data provider. Higher-risk datasets shall be referred to Legal, Compliance or the responsible AI governance function as appropriate.
AI Training Data Register
REACT shall maintain an AI Training Data Register providing traceability of datasets used for AI development. At minimum, the register should record:
- Dataset ID — Unique identifier
- Dataset name — Internal dataset name
- Source — Origin of data
- Data owner — REACT / member / third party
- Data type — Images, text, metadata, etc.
- Volume — Approximate number of records/files
- AI system — Model using the dataset
- Purpose — Training / testing / validation
- Permission / legal basis — Supporting basis
- Restrictions — Member, licence or use restrictions
- Personal data — Yes / No
- Security classification — Applicable classification
- Approval — Responsible approver
- Date approved — Approval date
- Retention — Retention period
- Status — Active / Archived / Deleted
Data Quality and Relevance
Training data should be relevant, sufficiently accurate and reasonably representative for the intended AI purpose. Where data is labelled for supervised learning, controls should reduce incorrect or inconsistent classifications. Where feasible, datasets should be evaluated for duplication, inappropriate content, significant bias and other quality issues that could materially affect model performance.
Security and Access
Training datasets shall be stored only in approved REACT environments. Access shall be limited to authorised personnel with a business need. Training data shall not be copied to personal devices, public AI services, unapproved cloud environments or other unauthorised systems. Appropriate logging, access control, backup and security measures shall be applied according to data sensitivity.
Third-Party AI Developers and Providers
Where REACT uses external developers, AI providers or technology partners, contractual and technical controls should prevent REACT or member data from being reused for unrelated purposes, incorporated into other customers' models, used to train general-purpose models, sold, commercially exploited or disclosed to unauthorised third parties unless explicitly approved.
Model Memorisation and Output Risk
Where appropriate, REACT shall assess whether trained models can reproduce protected, confidential or personal information contained in training datasets. Unacceptable memorisation or reproduction shall be mitigated before deployment.
Human Oversight
Training data and AI-generated classifications remain subject to appropriate human oversight. AI outputs used for enforcement decisions should be treated as decision-support information unless the relevant AI system has been specifically assessed and approved for another level of automation.
Retention and Deletion
Training data shall not be retained indefinitely without justification. Retention shall consider model-development requirements, member instructions, contractual and legal obligations, validation needs, and security and privacy risks. Deletion should be documented in the AI Training Data Register. Where withdrawal from an already trained model is technically impracticable, the matter shall be assessed and the data provider informed where appropriate.
Roles and Responsibilities
The AI System Owner is responsible for ensuring that only approved datasets are used. AI developers and analysts must follow this Policy and maintain dataset traceability. Information Security establishes security requirements. Legal and Data Protection provide specialist assessment where required. Management ensures appropriate governance and resources for responsible AI development.
Exceptions
Exceptions must be documented, risk-assessed and approved by the appropriate responsible function before the relevant data is used.