GPAI Ledger › GPAI Training Transparency tracker (AI Accountability Lab (AIAL)) › Capture 12 Sep 2026
Midjourney_Midjourney_2026_09_08 — capture 20260912T062140Z
Filed under AI Accountability Lab (AIAL) — GPAI Training Transparency tracker, the source this project captured it from; the document itself is the filing of the model named above.
| Provider | provider not identified by this project |
|---|---|
| Target | AIAL archived copy — https://raw.githubusercontent.com/AIAccountabilityLab/gpai-training-transparency/c878ebbe5a3d7585857ed203f59f34551d7e72f7/public/archive/Midjourney_Midjourney_2026_09_08.pdf |
| Fetched (UTC) | 2026-09-12T06:21:40Z |
| Upstream commit | 8 Sep 2026 — c878ebbe5a3d (when this state began to stand in the upstream repository; this archive fetched it at the time above, not then) |
| Stored file | 682681e354d89c8c0397fa56a52b13077ce1cddc152d9f43f46e9f6b55c3210d.pdf (412,020 bytes) |
| SHA-256 | 682681e354d89c8c0397fa56a52b13077ce1cddc152d9f43f46e9f6b55c3210d |
| OpenTimestamps proof | 682681e354d89c8c0397fa56a52b13077ce1cddc152d9f43f46e9f6b55c3210d.pdf.20260912T062140Z.ots (calendar-attested; anchored in bitcoin over time) |
| Wayback | not saved |
| Prior capture of this target | — first capture of this target |
| Notes | 16 character(s) (typographic ligatures such as the 'ffi' in 'Office') could not be decoded from the source file's embedded font and are shown as �; this affects only this text rendering — the stored file is exact |
Verify: sha256sum 682681e354d89c8c0397fa56a52b13077ce1cddc152d9f43f46e9f6b55c3210d.pdf must equal the hash above (the filename IS the expected hash); ots verify 682681e354d89c8c0397fa56a52b13077ce1cddc152d9f43f46e9f6b55c3210d.pdf.20260912T062140Z.ots -f 682681e354d89c8c0397fa56a52b13077ce1cddc152d9f43f46e9f6b55c3210d.pdf (opentimestamps.org) proves the bytes existed no later than the attestation time — an upper bound on the capture time; the fetch time above is the archive's own record (a freshly captured proof reports 'pending' here: the calendars anchor within hours, but this archive only upgrades the stored proof to its anchor on a later run, so expect a day or two). ots verify needs a local Bitcoin Core node (a pruned one is fine); without one, ots info on the proof prints the attesting block height and merkle path to check on any block explorer.
Extracted text
Machine-extracted text (layout may be lost; the authoritative content is the stored file above).
Midjourney / Documentation / Midjourney Policies Categories Version of Summary: Version #1 Last update: March 17, 2026 Provider name and contact details: Midjourney, Inc. Authorised representative name and contact details: Mark Foster (mark.foster@samadvisory.eu) Versioned model name(s): Midjourney Image and Video family of models Model dependencies: Midjourney V8, V8.1, and V8.2 Date of placement of the model on the Union market: March 17, 2026 ☒ Text☒ Less than 1 billion tokens □ 1 billion to 10 trillions The family of models trains on datasets containing publicly Ask AI Public Summary of Training Content – Midjourney https://docs.midjourney.com/hc/en-us/articles/4806708031... 1 of 7 08/09/2026, 17:08 tokens □ More than 10 trillions tokens available text annotations and image captions. ☒ Image □ Less than 1 million images □ 1 million to 1 billion images ☒ More than 1 billion images The family of models trains on datasets containing photography, visual art works, illustrations, textual metadata associated with images, human-provided annotations, prompts and preference data. □ Audio □ Less than 10 000 hours □ 10 000 to 1 million hours □ More than 1 million hours ☒ Video □ Less than 10 000 hours □ 10 000 to 1 million hours ☒ More than 1 million hours The family of models trains on datasets containing video clips, video effects, textual metadata associated with these videos, human-provided annotations, prompts and preference data. □ Other Latest date of data acquisition/ collection for model training: The data used to train the model includes datasets with varying cutoff dates. Datasets were used to train the models as late as March 2026. The model is continuously trained and may undergo additional �ne-tuning which may be released in new versions. Description of the linguistic characteristics of the overall training data: Training sources include both European and non-European languages. Other relevant characteristics of the overall training data: Midjourney training data represents a large- scale and diverse range of data including publicly-available websites, images, text, and video. The datasets are �ne tuned for Midjourney’s purposes. Midjourney training data undergoes several processing steps during training, including: deduplication, removal of low quality images, safety �ltering, privacy processing to �lter or remove sensitive personal information, and categorization based on relevance, quality, or image formats. Ask AI Public Summary of Training Content – Midjourney https://docs.midjourney.com/hc/en-us/articles/4806708031... 2 of 7 08/09/2026, 17:08 Additional comments (optional): N/A Have you used publicly available datasets to train the model? ☒ Yes □ No If yes, specify the modality(ies) of the content covered by the datasets concerned: ☒ Text ☒ Image □ Video □ Audio □ Other If so, please specify... List of large publicly available datasets: Midjourney datasets are composed of a wide variety of data publicly accessible online. General description of other publicly available datasets not listed above: Public datasets are �ltered and �ne tuned for quality and safety as described above and exclude sources that have opted out of training using web controls such as robots.txt �les. Additional comments (optional): N/A Have you concluded transactional commercial licensing agreement(s) with rightsholder(s) or with their representatives? ☒ Yes □ No If yes, specify the modality(ies) of the content covered by the datasets concerned: ☒ Text ☒ Image □ Video Have you obtained private ☒ Yes □ No Ask AI Public Summary of Training Content – Midjourney https://docs.midjourney.com/hc/en-us/articles/4806708031... 3 of 7 08/09/2026, 17:08 datasets from third parties that are not licensed as described in Section 2.2.1, such as data obtained from providers of private databases, or data intermediaries? If yes, specify the modality(ies) of the content covered by the datasets concerned: ☒ Text ☒ Image □ Video □ Audio □ Other If so, please specify... If publicly known, list private datasets obtained from other third parties: Some Midjourney datasets are purchased or licensed from third parties. These deals are bound by con�dentiality obligations. General description of non- publicly known private datasets obtained from third parties Data is covered by agreements that outline the party’s roles and responsibilities with respect to the datasets Midjourney uses. Additional comments (optional): N/A Were crawlers used by the provider or on behalf of? ☒ Yes □ No If yes, specify crawler name(s)/ identi�er(s): Deals are bound by con�dentiality obligations Purposes of the crawler(s): Crawlers are used to obtain publicly available content for training purposes. General description of crawler behaviour: Crawlers are used to �nd information, scan websites, and extract data from lawfully and publicly accessible resources. Crawlers are designed to respect robots.txt rules and do not circumvent technological measures. Period of data collection: 2023 to present Comprehensive description of the type of content and online sources crawled: Crawled data pulls from publicly available online material including images and associated text descriptions. The crawled data was �ltered for safety, quality, and privacy. Type of modality covered: ☒ Text ☒ Image ☒ Video □ Audio □ Other If so, please specify.. Ask AI Public Summary of Training Content – Midjourney https://docs.midjourney.com/hc/en-us/articles/4806708031... 4 of 7 08/09/2026, 17:08 Summary of the most relevant domain names crawled: The crawled data includes data from publicly available online websites and sources that encompass a large variety of content types and languages. Toplevel domain names crawled include: .com, .org, and .net as well as other global sites. Additional comments (optional): N/A Was data from user interactions with the AI model (e.g. user input and prompts) used to train the model? ☒ Yes □ No Was data collected from user interactions with the provider’s other services or products used to train the model? □ Yes ☒ No If yes, provide a general description of the provider’s services or products that were used to collect the user data: N/A Type of modality covered: ☒ Text ☒ Image □ Video □ Audio □ Other If so, please specify... Additional comments (optional): Midjourney adheres to its Privacy Policy and Terms of Service, as applicable, for user data training. Was synthetic AI-generated data created by the provider or on their behalf to train the model? ☒ Yes □ No If yes, modality of the synthetic data: ☒ Text ☒ Image □ Video □ Audio □ Other If so, please specify... Ask AI Public Summary of Training Content – Midjourney https://docs.midjourney.com/hc/en-us/articles/4806708031... 5 of 7 08/09/2026, 17:08 If yes, specify the general- purpose AI model(s) used to generate the synthetic data if available on the market: Midjourney may use internal models to generate synthetic data for training. Information about other AI models, including provider’s own AI model(s) not available on the market, used to generate synthetic data to train the model to which this Summary applies: N/A Additional comments (optional): N/A Have data sources other than those described in Sections 2.1 to 2.5 been used to train the model? ☒ Yes □ No If yes, provide a narrative description of these data sources and the data: Midjourney used datasets that it has acquired through its business operations. Additional comments (optional): N/A Are you a Signatory to the Code of Practice for general-purpose AI models that includes commitments to respect reservations of rights from the TDM exception or limitation? □ Yes ☒ No Ask AI Public Summary of Training Content – Midjourney https://docs.midjourney.com/hc/en-us/articles/4806708031... 6 of 7 08/09/2026, 17:08 Describe the measures implemented before model training to respect reservations of rights from the TDM exception or limitation before and during data collection, including the opt-out protocols and solutions honoured by the provider or, as applicable, by third parties from which datasets have been obtained: Midjourney implemented safety and quality �ltering, content moderation, honored robots.txt �les instructions where appropriate, and conducted privacy processing to �lter or remove sensitive personal information from training data. Additional comments (optional): N/A General description of measures taken: Midjourney adheres to applicable laws and best practices with respect to removing illegal content. Midjourney training data undergoes safety �ltering to remove certain data with known risk of containing child sexual abuse material (CSAM) and other categories of sensitive or disallowed content. See also rules of conduct for its users at: https:// docs.midjourney.com/hc/en-us/ articles/32013696484109-Community- Guidelines. Other relevant information about data processing (optional): N/A Midjourney Website Midjourney Discord Server Ask AI Public Summary of Training Content – Midjourney https://docs.midjourney.com/hc/en-us/articles/4806708031... 7 of 7 08/09/2026, 17:08