GPAI Ledger The public record of EU AI Act training-data summaries

GPAI LedgerMuse Spark (Meta) › Capture 20 Aug 2026

Muse Spark — capture 20260820T110728Z

ProviderMeta
Targetprovider site — https://scontent-mxp1-1.xx.fbcdn.net/m1/v/t0.84174-6/An_FRNs4SBgBP-25psN50ZmPdjiJKewbOP2ZVfOXrTF_PfKb16y6pkNhkYa2ekzAfEb8rqL1RBf0KMJ4eGnMhW-UykoM22GMFn4ZSJPvk8NsFeWFVzFOCQCqQU8gZiQNBv59Bxs?_nc_gid=…&_nc_oc=…&ccb=…&oh=…&oe=…&_nc_sid=… (signed URL; token masked, not linked)
Fetched (UTC)2026-08-20T11:07:28Z
Stored filefb8e4a1daaba1538587bee5658f0afa2b52b082edfcc3132b5d215e3c2542217.pdf (257,787 bytes)
SHA-256fb8e4a1daaba1538587bee5658f0afa2b52b082edfcc3132b5d215e3c2542217
OpenTimestamps prooffb8e4a1daaba1538587bee5658f0afa2b52b082edfcc3132b5d215e3c2542217.pdf.ots (calendar-attested; anchored in bitcoin over time)
Waybacknot saved
Prior capture of this target— first capture of this target

Verify: sha256sum fb8e4a1daaba1538587bee5658f0afa2b52b082edfcc3132b5d215e3c2542217.pdf must equal the hash above (the filename IS the expected hash); ots verify fb8e4a1daaba1538587bee5658f0afa2b52b082edfcc3132b5d215e3c2542217.pdf.ots -f fb8e4a1daaba1538587bee5658f0afa2b52b082edfcc3132b5d215e3c2542217.pdf (opentimestamps.org) proves the capture time (fresh proofs report 'pending' until bitcoin-anchored, typically within a day).

Extracted text

Machine-extracted text (layout may be lost; the authoritative content is the stored file above).

Public Summary of Training ContentVersion of the Summary: Document version number: V3
Last Update: August 4, 20261. General information
1.1 Provider Identification
Provider's name and contact details: Meta Platforms Ireland Limited (MPIL)Authorized representative name and contact details: Support Contact: https://ai.meta.com/help1.2. Model identificationVersioned model name(s): Muse SparkModel dependencies: N/ADate of placement of the model on the Union market:April 8, 20261.3. Modalities, overall training data size and other characteristics
ModalitySelect the modalities present in the training data, to the extent that they are identifiable
Training data sizeFor each selected modality, select the range within which the estimated total training data size for that modality falls. Dynamic datasets may beexcluded from the estimation.
Types of contentFor each selected modality, provide a general description of the type of content that has been included in the training data
Text
More than 10 trillions tokens
This dataset comprises multimodal content sourced from publicly available data, data provided by third parties and information from Meta’s products and services, curated and enriched by external vendor networks and Meta personnel.Perception (Image & Video)
More than 10 trillions tokens
This dataset comprises multimodal content sourced from publicly available data, data provided by third parties and information from Meta’s products and services, curated and enriched by external vendor networks and Meta personnel.Audio (Excluding audio that is part of video, as this should be reported under the “video” modality instead. Furthermore, the Commissionunderstands the modality of ‘audio’ to include ‘speech’.)
More than 10 trillions tokens
This dataset comprises multimodal content sourced from publicly available data, data provided by third parties and information from Meta’s products and services, curated and enriched by external vendor networks and Meta personnel.Other
 N/A1.3.1 Other CharacteristicsLatest date of data: acquisition/collection for model training: Up to July 2026Description of the linguistic characteristics of the overall training data: Multiple languages and geographies, including EU official languages.Other relevant characteristics of the overall training data: N/AAdditional comments (optional): For more about the information we use for AI at Meta, where it comes from and how it works, see: https://transparency.meta.com/features/ai-at-meta-training-data/                              2. List of data sources 2.1. Publicly available datasetsHave you used publicly available datasets to train the model?YesIf yes, specify the modality(ies) of the content covered by the datasets concerned: Text Image Video Audio Other (Please Specify):
List of large publicly available datasets: Training Data: A mix of publicly available, licensed data and information from Meta’s products and services, including publicly shared posts from Instagram and Facebook. Learn more in our Privacy Center.
General description of other publicly available datasets not listed above:The model has been trained on multi-modal and multilingual (including EU language) data from a range of publicly available datasets.Additional comments (optional): N/A2.2. Private non-publicly available datasets obtained from third parties2.2.1. Datasets commercially licensed by rightsholders or their representativesHave you concluded transactional commercial licensing agreement(s) with rightsholder(s) or with their representatives? YesIf yes, specify the modality(ies) of the content covered by the datasets concerned: Text Image Video Audio Other (Please Specify):
2.2.2. Private datasets obtained from other third partiesHave you obtained private datasets from third parties that are notlicensed as described in Section 2.2.1, such as data obtained fromproviders of private databases, or data intermediaries?YesIf yes, specify the modality(ies) of the content covered by the datasetsconcerned: Text Image Video Audio Other (Please Specify):
If publicly known, list private datasets obtained from other third parties: N/A
General description of non-publicly known private datasets obtained from third parties:
This dataset comprises multimodal content sourced from publicly available data, data provided by third parties and information from Meta’s products and services, curated and enriched by external vendor networks and Meta personnel.Additional comments (optional): N/A2.3. Data crawled and scraped from online sourcesWere crawlers used by the provider or on behalf of?YesIf yes, specify crawler name(s)/identifier(s):Purposes of the crawler(s) It is Meta’s policy to provide information on Meta’s web crawlers through Meta’s developer center (https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/). General description of crawlerbehaviour: Meta crawlers are designed to respect standard web protocols and ethical data collection practices.Period of data collection: Up to July 2026

Public Summary of Training Content
Comprehensive description of the type of content and online sources crawled: See the developer center resource linked above for more information about Meta's web crawlers.                                                                Type of modality covered: Text Image Video Audio Other (Please Specify): Summary of the most relevant domainnames crawled: See the developer center resource linked above for more information about Meta's web crawlers. Additional comments (optional): N/A2.4. User dataWas data from user interactions with the AI model (e.g. user input and prompts) used to train the model? YesWas data collected from user interactions with the provider’s other services or products used to train the model? YesIf yes, provide a general description of the provider’s services or products that were used to collect the user data:Data from 1st party services like Facebook and Instagram.Type of modality covered Text Image Video Audio Other (Please Specify): Additional comments (optional): N/A2.5. Synthetic dataWas synthetic AI-generated data created by the provider or on their behalf to train the model? YesIf yes, modality of the synthetic data: Text Image Video Audio Other (Please Specify):
If yes, specify the general-purpose AI model(s) used to generate the synthetic data if available on the market:
Synthetic data used in training the model was generated by a collection of generative AI models, to generate training data such as ideal responses, golden training datasets and high quality captions for images. For more information on how Meta uses synthetic data for AI at Meta see: https://transparency.meta.com/features/ai-at-meta-training-data/Information about other AI models, including provider’s own AI model(s) not available on the market, used to generate synthetic data to train the model to which this Summary applies: See above Additional comments (optional): N/A2.6. Other sources of dataHave data sources other than those described in Sections 2.1 to 2.5 been used to train the model? NoIf yes, provide a narrative description of these data sources and the dataN/AAdditional comments (optional): N/A3. Data processing aspects3.1. Respect of reservation of rights from text and data mining exception or limitationAre you a Signatory to the Code of Practice for general purpose AI models that includes commitments to respect reservations of rights from the TDM exception or limitation? NoDescribe the measures implemented before model training to respect reservations of rights from the TDM exception or limitation before and during data collection, including the opt-out protocols and solutions honoured by the provider or, as applicable, by third parties from which datasets have been obtained
It is Meta’s policy to provide information on Meta’s web crawlers through Meta’s developer center (https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/). Meta’s developer page and information relating to crawlers employed by Meta may be updated from time to time.Additional comments (optional): N/A3.2. Removal of illegal content
General description of measures taken:
Data used for training was subject to several curation methodologies such as cleaning, filtering, summarization and ratings. Meta takes proactive measures aimed at preventing the inclusion of illegal content in generative AI training datasets.3.2 Other information (optional)Other relevant information about data processing (optional):N/A