GPAI Ledger The public record of EU AI Act training-data summaries

GPAI LedgerMuse Spark (Meta) › Capture 11 Aug 2026

Muse Spark — capture 20260811T111428Z

ProviderMeta
TargetAIAL archived copy — https://aial.ie/research/gpai-training-transparency/archive/Muse_Spark_2026_07_21.pdf
Fetched (UTC)2026-08-11T11:14:28Z
Stored file6d0dce769e8b2edd5ef6f6963fb31facca5211c9344970016434cf300d2a3729.pdf (258,200 bytes)
SHA-2566d0dce769e8b2edd5ef6f6963fb31facca5211c9344970016434cf300d2a3729
OpenTimestamps proof6d0dce769e8b2edd5ef6f6963fb31facca5211c9344970016434cf300d2a3729.pdf.ots (calendar-attested; anchored in bitcoin over time)
Waybacknot saved
Prior capture of this target— first capture of this target
Notestext_sha256 recorded 20 Aug 2026 from the extracted.txt stored at capture time (bootstrap captures predate this field); raw bytes unchanged

Verify: sha256sum 6d0dce769e8b2edd5ef6f6963fb31facca5211c9344970016434cf300d2a3729.pdf must equal the hash above (the filename IS the expected hash); ots verify 6d0dce769e8b2edd5ef6f6963fb31facca5211c9344970016434cf300d2a3729.pdf.ots -f 6d0dce769e8b2edd5ef6f6963fb31facca5211c9344970016434cf300d2a3729.pdf (opentimestamps.org) proves the capture time (fresh proofs report 'pending' until bitcoin-anchored, typically within a day).

Extracted text

Machine-extracted text (layout may be lost; the authoritative content is the stored file above).

Public Summary of Training Content
Version of the Summary: Document version number: V2
Last Update: 7/8/261. General information
1.1 Provider Identification
Provider's name and contact details: Meta Platforms Ireland Limited (MPIL)Authorized representative name and contact details: Support Contact: https://ai.meta.com/help1.2. Model identificationVersioned model name(s): Muse SparkModel dependencies: N/A
Date of placement of the model on the Union market:4/8/261.3. Modalities, overall training data size and other characteristics
ModalitySelect the modalities present in the training data, to the extent that they are identifiable
Training data sizeFor each selected modality, select the range within which the estimated total training data size for that modality falls. Dynamic datasets may beexcluded from the estimation.
Types of contentFor each selected modality, provide a general description of the type of content that has been included in the training data
Text
More than 10 trillions tokens
This dataset comprises multimodal content sourced from publicly available data, data provided by third parties and information from Meta’s products and services, curated and enriched by external vendor networks and Meta personnel.Perception (Image & Video)
More than 10 trillions tokens
This dataset comprises multimodal content sourced from publicly available data, data provided by third parties and information from Meta’s products and services, curated and enriched by external vendor networks and Meta personnel.Audio (Excluding audio that is part of video, as this should be reported under the “video” modality instead. Furthermore, the Commissionunderstands the modality of ‘audio’ to include ‘speech’.)
More than 10 trillions tokens
This dataset comprises multimodal content sourced from publicly available data, data provided by third parties and information from Meta’s products and services, curated and enriched by external vendor networks and Meta personnel.Other
 N/A1.3.1 Other CharacteristicsLatest date of data: acquisition/collection for model training: Up to June 2026Description of the linguistic characteristics of the overall training data: Multiple languages and geographies, including EU official languages.Other relevant characteristics of the overall training data: N/AAdditional comments(optional):
For more about the information we use for AI at Meta, where it comes from and how it works, see: https://transparency.meta.com/features/ai-at-meta-training-data/                              2. List of data sources 2.1. Publicly available datasetsHave you used publicly available datasets to train the model?YesIf yes, specify the modality(ies) of the content covered by the datasets concerned: Text Image Video Audio Other (Please Specify):
List of large publicly available datasets: Training Data: A mix of publicly available, licensed data and information from Meta’s products and services, including publicly shared posts from Instagram and Facebook. Learn more in our Privacy Center.
General description of other publicly available datasets not listed above:The model has been trained on multi-modal and multi-lingual (including EU language) data from a range of publicly available datasets.Additional comments (optional): N/A2.2. Private non-publicly available datasets obtained from third parties2.2.1. Datasets commercially licensed by rightsholders or their representativesHave you concluded transactional commercial licensing agreement(s) with rightsholder(s) or with their representatives? YesIf yes, specify the modality(ies) of the content covered by the datasets concerned: Text Image Video Audio Other (Please Specify):
2.2.2. Private datasets obtained from other third partiesYesHave you obtained private datasets from third parties that are not licensed as described in Section 2.2.1, such as data obtained from providers of private databases, or data intermediaries? Text Image Video Audio Other (Please Specify):
If publicly known, list private datasets obtained from other third parties: N/A
General description of non-publicly known private datasets obtained from third parties:
This dataset comprises multimodal content sourced from publicly available data, data provided by third parties and information from Meta’s products and services, curated and enriched by external vendor networks and Meta personnel.Additional comments (optional): N/A2.3. Data crawled and scraped from online sourcesWere crawlers used by the provider or on behalf of?Yes
If yes, specify crawler name(s)/identifier(s):Purposes of the crawler(s) It is Meta’s policy to provide information on Meta’s web crawlers through Meta’s developer center (https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/). General description of crawlerbehaviour: Meta crawlers are designed to respect standard web protocols and ethical data collection practices.Period of data collection: Up to June 2026

Public Summary of Training Content
Comprehensive description of the type of content and online sources crawled: See the developer center resource linked above for more information about Meta's web crawlers.                                                                Type of modality covered: Text Image Video Audio Other (Please Specify): Summary of the most relevant domainnames crawled: See the developer center resource linked above for more information about Meta's web crawlers. Additional comments (optional): N/A2.4. User dataWas data from user interactions with the AI model (e.g. user input and prompts) used to train the model? YesWas data collected from user interactions with the provider’s other services or products used to train the model? YesIf yes, provide a general description of the provider’s services or products that were used to collect the user data:Data from 1st party services like Facebook and Instagram.Type of modality covered Text Image Video Audio Other (Please Specify): Additional comments (optional): N/A2.5. Synthetic dataWas synthetic AI-generated data created by the provider or on their behalf to train the model? YesIf yes, modality of the synthetic data: Text Image Video Audio Other (Please Specify):
If yes, specify the general-purpose AI model(s) used to generate the synthetic data if available on the market:
Synthetic data used in training the model was generated by a collection of generative AI models, to generate training data such as ideal responses, golden training datasets and high quality captions for images. For more information on how Meta uses synthetic data for AI at Meta see: https://transparency.meta.com/features/ai-at-meta-training-data/Information about other AI models, including provider’s own AI model(s) not available on the market, used to generate synthetic data to train the model to which this Summary applies: See above Additional comments (optional): N/A2.6. Other sources of dataHave data sources other than those described in Sections 2.1 to 2.5 been used to train the model? NoIf yes, provide a narrative description of these data sources and the dataN/AAdditional comments (optional): N/A3. Data processing aspects3.1. Respect of reservation of rights from text and data mining exception or limitationAre you a Signatory to the Code of Practice for general purpose AI models that includes commitments to respect reservations of rights from the TDM exception or limitation? NoDescribe the measures implemented before model training to respect reservations of rights from the TDM exception or limitation before and during data collection, including the opt-out protocols and solutions honoured by the provider or, as applicable, by third parties from which datasets have been obtained
It is Meta’s policy to provide information on Meta’s web crawlers through Meta’s developer center (https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/). Meta’s developer page and information relating to crawlers employed by Meta may be updated from time to time.Additional comments (optional): N/A3.2. Removal of illegal content
General description of measures taken:
Data used for training was subject to several curation methodologies such as cleaning, filtering, summarization and ratings. Meta takes proactive measures aimed at preventing the inclusion of illegal content in generative AI training datasets.3.2 Other information (optional)Other relevant information about data processing (optional):N/A