GPAI Ledger The public record of EU AI Act training-data summaries

GPAI LedgerGPAI Training Transparency tracker (AI Accountability Lab (AIAL)) › Capture 12 Sep 2026

Minimax_H3_2026_09_11 — capture 20260912T062833Z

Filed under AI Accountability Lab (AIAL) — GPAI Training Transparency tracker, the source this project captured it from; the document itself is the filing of the model named above.

Providerprovider not identified by this project
TargetAIAL archived copy — https://raw.githubusercontent.com/AIAccountabilityLab/gpai-training-transparency/a030eede7e4582b27dd04a8453647824d8da9784/public/archive/Minimax_H3_2026_09_11.pdf
Fetched (UTC)2026-09-12T06:28:33Z
Upstream commit11 Sep 2026 — a030eede7e45 (when this state began to stand in the upstream repository; this archive fetched it at the time above, not then)
Stored filec194fd517d43b7e34635a0ec90d785d32cae9f33e70cddc28f3d885a8362365e.pdf (427,523 bytes)
SHA-256c194fd517d43b7e34635a0ec90d785d32cae9f33e70cddc28f3d885a8362365e
OpenTimestamps proofc194fd517d43b7e34635a0ec90d785d32cae9f33e70cddc28f3d885a8362365e.pdf.20260912T062833Z.ots (calendar-attested; anchored in bitcoin over time)
Waybacknot saved
Prior capture of this target— first capture of this target

Verify: sha256sum c194fd517d43b7e34635a0ec90d785d32cae9f33e70cddc28f3d885a8362365e.pdf must equal the hash above (the filename IS the expected hash); ots verify c194fd517d43b7e34635a0ec90d785d32cae9f33e70cddc28f3d885a8362365e.pdf.20260912T062833Z.ots -f c194fd517d43b7e34635a0ec90d785d32cae9f33e70cddc28f3d885a8362365e.pdf (opentimestamps.org) proves the bytes existed no later than the attestation time — an upper bound on the capture time; the fetch time above is the archive's own record (a freshly captured proof reports 'pending' here: the calendars anchor within hours, but this archive only upgrades the stored proof to its anchor on a later run, so expect a day or two). ots verify needs a local Bitcoin Core node (a pruned one is fine); without one, ots info on the proof prints the attesting block height and merkle path to check on any block explorer.

Extracted text

Machine-extracted text (layout may be lost; the authoritative content is the stored file above).

1
 Public Summary of Training Content for MiniMax H3 This template is provided by the European Commission and required to be filled in by providers of general-purpose AI models prior to their placing on the Union market in order to comply with their obligation under Article 53 (1)(d) of Regulation (EU) 2024/1689 (AI Act).  For more information and guidance see Commission’s Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models | Shaping Europe’s digital future.  Version of the Summary:  1.0 Last update:  08/02/2026 General information 1. General information 1.1. Provider identification  Provider name and contact details:  Nanonoble Pte. Ltd. Authorised representative name and contact details: Prighter GmbH  1.2. Model identification  Versioned model name(s): MiniMax H3 Model dependencies: N/A Date of placement of the model on the Union market:  08/02/2026   1.3 Modalities, overall training data size and other characteristics This Section requires general information about the overall training data after pre-processing and before the training of the model.   Modality Select the modalities present in the training data, to the extent that they are identifiable
Training data size For each selected modality, select the range within which the estimated total training data size for that modality falls. Dynamic datasets may be excluded from the estimation.
Types of content For each selected modality, provide a general description of the type of content that has been included in the training data.
 ☒ Text
☐ Less than 1 billion tokens ☒ 1billion to 10 trillions tokens ☐ More than 10 trillions tokens  Alternatively, specify the approximate size in a different measurement unit:  N/A
 Text data spanning broad domains and multimodal descriptions, including textual descriptions of images and videos.
☒ Image ☐ Less than 1 million images ☐ 1Million to1 billion images ☒ More than 1 billion images A large-scale, diverse collection of images spanning broad domains and scenarios.

2
☒ Audio1 ☐ Less than 10 000 hours ☐ 10 000 to1 million hours  ☒ More than 1 million hours Audio data spanning diverse sound types and acoustic scenarios.
☒ Video  ☐ Less than 10 000 hours ☐ 10 000 to1 million hours  ☒ More than 1 million hours A large-scale, diverse collection of video clips spanning broad domains and scenarios.
☐ Other Specify the modality and for each one indicate approximate size and unit of measurement
 Latest date of data acquisition/collection for model training: July 2026 Description of the linguistic characteristics of the overall training data:  The training data covers a diverse range of global languages. No specific geographic region was intentionally excluded from the data collection process. Other relevant characteristics of the overall training data: The training data covers text, image, video and audio modalities across broad domains and scenarios. Additional comments (optional):   2. List of data sources
2. List of data sources This Section requires information about specific sources of data used to train the general-purpose AI model. In this section “dataset” should be understood as a single, pre-packaged collection of data. The filtering and pre-processing of data collected from the same pre-packaged collection should not be considered a new dataset to be disclosed separately in the sections below. If a particular dataset can be assigned to more than one of the categories below, providers should select the most relevant category and only report the dataset in that category, except in the case of synthetic data (see Section 2.5).    2.1. Publicly available datasets    This Section requires information about datasets that were used to train the model and which have been compiled by a third party, are made available publicly for free, and are readily downloadable as a whole or in predefined chunks, such as datasets and collections available on public repositories and online platforms, specialised websites, or snapshots of common crawl. The public availability of the datasets for free does not mean that the content at issue is necessarily free of rights since it may be subject to licensing arrangements or conditions of use (e.g., certain free and/or open licenses may determine the scope of the uses, including prohibiting uses relating to model training).  A dataset is considered to be “large” if the total data size for any one of the modalities contained in the dataset exceeds 3% of the size of all publicly available datasets for that modality used for training. The size of the dataset should be based on its size after pre-processing (for example filtering), and without splitting the dataset to prevent reporting circumvention.      Have you used publicly available datasets to train the model?   ☒ Yes     ☐ No If yes, specify the modality(ies) of the content covered by the datasets concerned: ☐ Text   ☒ Image   ☒ Video  ☐ Audio ☐ Other  If so, please specify…  1 Excluding audio that is part of video, as this should be reported under the “video” modality instead. Furthermore, the Commission understands the modality of ‘audio’ to include ‘speech’.

3
List of large publicly available datasets:  The training data comprises a mix of publicly available and licensed data. This may include image and video datasets made available by third parties through public repositories and online platforms. General description of other publicly available datasets not listed above: Publicly available datasets covering image and video modalities.
Additional comments (optional):   2.2 Private non-publicly available datasets obtained from third parties This Section requires information about private non-publicly available datasets of third parties that are not publicly available and not disclosed under Section 2.1. These include: 1) datasets for which transactional commercial licensing agreements were concluded between the provider and the rightsholders or their representatives, including by collective management organisations and legitimate content aggregators who have the right to collectively license works on behalf of rightsholders (Section 2.2.1); 2) other private datasets obtained through data intermediaries, non-publicly available databases and datasets of third parties for which transactional commercial licenses have not been concluded with rightsholders or their representatives (Section 2.2.2).   2.2.1. Datasets commercially licensed by rightsholders or their representatives Have you concluded transactional commercial licensing agreement(s) with rightsholder(s) or with their representatives? ☒ Yes     ☐ No If yes, specify the modality(ies) of the content covered by the datasets concerned:  ☐ Text   ☒ Image   ☒ Video   ☒ Audio ☐ Other  If so, please specify… 2.2.2. Private datasets obtained from other third parties Have you obtained private datasets from third parties that are not licensed as described in Section 2.2.1, such as data obtained from providers of private databases, or data intermediaries? ☒ Yes     ☐ No
If yes, specify the modality(ies) of the content covered by the datasets concerned:  ☐ Text   ☒ Image   ☒ Video  ☒ Audio ☐ Other If so, please specify… If publicly known, list private datasets obtained from other third parties:  N/A General description of non-publicly known private datasets obtained from third parties  Datasets have been obtained from third parties on a licensed basis. Such datasets comprise content across image, video and audio modalities. Additional comments (optional):   2.3 Data crawled and scraped from online sources This Section requires information about crawled, scraped data, or otherwise compiled from online sources directly by the provider of the model or on their behalf (i.e. excluding publicly available datasets already compiled by third parties and made available on platforms such as common crawl that are covered under Section 2.1).

4
Were crawlers used by the provider or on behalf of? ☒ Yes     ☐ No If yes, specify crawler name(s)/identifier(s): MINIMAX_UA Purposes of the crawler(s): MINIMAX_UA is used to collect publicly accessible image, video and audio content that may be used for model training, testing and validation. General description of crawler behaviour:  MINIMAX_UA respects robots.txt instructions and does not circumvent paywalls, CAPTCHAs, login pages or password-protected content. Period of data collection: Up to July 2026 Comprehensive description of the type of content and online sources crawled: Publicly accessible online content, primarily images, videos and audio. Type of modality covered:  ☐ Text   ☒ Image   ☒ Video  ☒ Audio  ☐ Other  If so, please specify… Summary of the most relevant domain names crawled: The relevant online sources span a broad range of domains and content contexts. Additional comments (optional):  2.4 User data This Section requires information about user data collected by all services and products of the provider, including through mail services, social media platforms, content platforms or interaction with the providers’  AI models and/or systems. This does not cover data licensed by users based on commercial transactional agreements described in Section 2.2.1., or customer data to fine-tune models for specific purposes.  Was data from user interactions with the AI model (e.g. user input and prompts) used to train the model?  ☐ Yes     ☒ No
Was data collected from user interactions with the provider’s other services or products used to train the model?   ☐ Yes     ☒ No If yes, provide a general description of the provider’s services or products that were used to collect the user data:  N/A Type of modality covered: N/A Additional comments (optional):   2.5 Synthetic data This Section requires information about synthetic data created by or on behalf of the provider for training the model directly on the outputs of another AI model, in particular through model distillation or model alignment (e.g. AI feedback through reinforcement learning). This does not include the use of AI models to clean or enrich data (e.g. AI-generated metadata to enrich or modify a dataset, such as creating depth maps or text descriptions of images). In case this concerns publicly available datasets as described in Section 2.1, these should be reported in that Section of the Template. In case this concerns synthetic datasets created by third parties on behalf of the provider , these should be reported in this Section of the Template instead of in Section 2.2.2.

5
Was synthetic AI-generated data created by the provider or on their behalf to train the model?    ☒ Yes     ☐ No If yes, modality of the synthetic data:  ☒ Text   ☐ Image   ☐ Video  ☐ Audio       ☐ Other  If so, please specify…  If yes, specify the general-purpose AI model(s) used to generate the synthetic data if available on the market: Synthetic text data, including textual descriptions of images, videos and audio, was generated using a range of general-purpose AI models, including multimodal understanding models and large language models.  Information about other AI models, including provider’s own AI model(s) not available on the market, used to generate synthetic data to train the model to which this Summary applies:
  Fine-tuned versions of internal models may be used to generate synthetic data. Additional comments (optional):    2.6 Other sources of data This Section requires information about data that does not fall under any of the categories in the previous Sections, for example data collected from offline sources, self-digitised media (e.g., digitised analog text context, images), datasets labelled by humans commissioned by the provider , or human generated data through reinforcement learning.   Have data sources other than those described in Sections 2.1 to 2.5 been used to train the model?  ☐ Yes     ☒ No If yes, provide a narrative description of these data sources and the data:    N/A Additional comments (optional):   1. Data processing aspects  3. Data processing aspects 3.1. Respect of reservation of rights from text and data mining exception or limitation  This Section concerns measures implemented by the provider to identify and comply with the reservation of rights from the text and data mining (TDM) exception or limitation expressed pursuant to Article 4(3) of Directive (EU) 2019/790, as outlined in the copyright policy put in place by the provider in accordance with Article 53(1)(c) AI Act.    Are you a Signatory to the Code of Practice for general-purpose AI models that includes commitments to respect reservations of rights from the TDM exception or limitation?
 ☐ Yes     ☒ No  Describe the measures implemented before model training to respect reservations of rights from the TDM exception or limitation before and during data collection, including the opt-out protocols and solutions honoured by the provider or, as applicable, by third parties from which datasets have been obtained:
  MINIMAX_UA respects robots.txt instructions and does not circumvent paywalls, CAPTCHAs, login pages or password-protected content.

6
Additional comments (optional):   3.2 Removal of illegal content This Section concerns measures taken to avoid or remove illegal content under Union law from the training data (such as blacklists, keywords, and model-based classifiers), without requiring disclosure of specific details about the provider’ s internal business practices or trade secrets. Such measures are advisable if the training data is likely to include illegal or unlawful content under Union law, in particular child sexual abuse material and terrorist content and the non-authorised use of material protected by intellectual property rights. Such measures do not include data selection practices, for example to increase the capability of the model.   General description of measures taken: The training data undergoes preprocessing and filtering, including the use of filtering models to exclude unsafe or harmful content.  3.3. Other information (optional) Other relevant information about data processing (optional):