GPAI Ledger The public record of EU AI Act training-data summaries

GPAI LedgerGPAI Training Transparency tracker (AI Accountability Lab (AIAL)) › Capture 12 Sep 2026

Midjourney_Midjourney_2026_09_08 — capture 20260912T062140Z

Filed under AI Accountability Lab (AIAL) — GPAI Training Transparency tracker, the source this project captured it from; the document itself is the filing of the model named above.

Providerprovider not identified by this project
TargetAIAL archived copy — https://raw.githubusercontent.com/AIAccountabilityLab/gpai-training-transparency/c878ebbe5a3d7585857ed203f59f34551d7e72f7/public/archive/Midjourney_Midjourney_2026_09_08.pdf
Fetched (UTC)2026-09-12T06:21:40Z
Upstream commit8 Sep 2026 — c878ebbe5a3d (when this state began to stand in the upstream repository; this archive fetched it at the time above, not then)
Stored file682681e354d89c8c0397fa56a52b13077ce1cddc152d9f43f46e9f6b55c3210d.pdf (412,020 bytes)
SHA-256682681e354d89c8c0397fa56a52b13077ce1cddc152d9f43f46e9f6b55c3210d
OpenTimestamps proof682681e354d89c8c0397fa56a52b13077ce1cddc152d9f43f46e9f6b55c3210d.pdf.20260912T062140Z.ots (calendar-attested; anchored in bitcoin over time)
Waybacknot saved
Prior capture of this target— first capture of this target
Notes16 character(s) (typographic ligatures such as the 'ffi' in 'Office') could not be decoded from the source file's embedded font and are shown as �; this affects only this text rendering — the stored file is exact

Verify: sha256sum 682681e354d89c8c0397fa56a52b13077ce1cddc152d9f43f46e9f6b55c3210d.pdf must equal the hash above (the filename IS the expected hash); ots verify 682681e354d89c8c0397fa56a52b13077ce1cddc152d9f43f46e9f6b55c3210d.pdf.20260912T062140Z.ots -f 682681e354d89c8c0397fa56a52b13077ce1cddc152d9f43f46e9f6b55c3210d.pdf (opentimestamps.org) proves the bytes existed no later than the attestation time — an upper bound on the capture time; the fetch time above is the archive's own record (a freshly captured proof reports 'pending' here: the calendars anchor within hours, but this archive only upgrades the stored proof to its anchor on a later run, so expect a day or two). ots verify needs a local Bitcoin Core node (a pruned one is fine); without one, ots info on the proof prints the attesting block height and merkle path to check on any block explorer.

Extracted text

Machine-extracted text (layout may be lost; the authoritative content is the stored file above).

Midjourney /  Documentation /  Midjourney Policies
Categories
Version of Summary:  Version #1
Last update:  March 17, 2026
Provider name and contact
details:   Midjourney, Inc.
Authorised representative
name and contact details:   Mark Foster (mark.foster@samadvisory.eu)
Versioned model name(s):  Midjourney Image and Video family of models
Model dependencies:  Midjourney V8, V8.1, and V8.2
Date of placement of the
model on the Union market:   March 17, 2026
☒ Text☒ Less than 1 billion tokens
□ 1 billion to 10 trillions
The family of models trains on
datasets containing publicly
Ask AI
Public Summary of Training Content – Midjourney https://docs.midjourney.com/hc/en-us/articles/4806708031...
1 of 7 08/09/2026, 17:08
tokens
□ More than 10 trillions
tokens
available text annotations and
image captions.
☒ Image
□ Less than 1 million images
□ 1 million to 1 billion images
☒ More than 1 billion images
The family of models trains on
datasets containing photography,
visual art works, illustrations,
textual metadata associated with
images, human-provided
annotations, prompts and
preference data.
□ Audio
□ Less than 10 000 hours
□ 10 000 to 1 million hours
□ More than 1 million hours

☒ Video
□ Less than 10 000 hours
□ 10 000 to 1 million hours
☒ More than 1 million hours
The family of models trains on
datasets containing video clips,
video effects, textual metadata
associated with these videos,
human-provided annotations,
prompts and preference data.
□ Other
Latest date of data acquisition/
collection for model training:
The data used to train the model includes
datasets with varying cutoff dates. Datasets
were used to train the models as late as March
2026. The model is continuously trained and
may undergo additional �ne-tuning which may
be released in new versions.
Description of the linguistic
characteristics of the overall
training data:
  Training sources include both European and
non-European languages.
Other relevant characteristics
of the overall training data:
Midjourney training data represents a large-
scale and diverse range of data including
publicly-available websites, images, text, and
video. The datasets are �ne tuned for
Midjourney’s purposes. Midjourney training data
undergoes several processing steps during
training, including: deduplication, removal of low
quality images, safety �ltering, privacy
processing to �lter or remove sensitive personal
information, and categorization based on
relevance, quality, or image formats.
Ask AI
Public Summary of Training Content – Midjourney https://docs.midjourney.com/hc/en-us/articles/4806708031...
2 of 7 08/09/2026, 17:08
Additional comments
(optional):   N/A
Have you used publicly
available datasets to train the
model?
 ☒ Yes   □ No
If yes, specify the modality(ies)
of the content covered by the
datasets concerned:
  ☒ Text   ☒ Image   □ Video   □ Audio
□ Other If so, please specify...
List of large publicly available
datasets:   Midjourney datasets are composed of a wide
variety of data publicly accessible online.
General description of other
publicly available datasets not
listed above:

Public datasets are �ltered and �ne tuned for
quality and safety as described above and
exclude sources that have opted out of training
using web controls such as robots.txt �les.
Additional comments
(optional):   N/A
Have you concluded
transactional commercial
licensing agreement(s) with
rightsholder(s) or with their
representatives?
 ☒ Yes   □ No
If yes, specify the modality(ies)
of the content covered by the
datasets concerned:
 ☒ Text   ☒ Image   □ Video
Have you obtained private ☒ Yes   □ No
Ask AI
Public Summary of Training Content – Midjourney https://docs.midjourney.com/hc/en-us/articles/4806708031...
3 of 7 08/09/2026, 17:08
datasets from third parties that
are not licensed as described in
Section 2.2.1, such as data
obtained from providers of
private databases, or data
intermediaries?
If yes, specify the modality(ies)
of the content covered by the
datasets concerned:
  ☒ Text   ☒ Image   □ Video   □ Audio
□ Other If so, please specify...
If publicly known, list private
datasets obtained from other
third parties:

Some Midjourney datasets are purchased or
licensed from third parties. These deals are
bound by con�dentiality obligations.
General description of non-
publicly known private datasets
obtained from third parties

Data is covered by agreements that outline the
party’s roles and responsibilities with respect to
the datasets Midjourney uses.
Additional comments (optional):  N/A
Were crawlers used by the
provider or on behalf of?  ☒ Yes   □ No
If yes, specify crawler name(s)/
identi�er(s):   Deals are bound by con�dentiality obligations
Purposes of the crawler(s):  Crawlers are used to obtain publicly available
content for training purposes.
General description of crawler
behaviour:
Crawlers are used to �nd information, scan
websites, and extract data from lawfully and
publicly accessible resources. Crawlers are
designed to respect robots.txt rules and do not
circumvent technological measures.
Period of data collection:  2023 to present
Comprehensive description of
the type of content and online
sources crawled:

Crawled data pulls from publicly available online
material including images and associated text
descriptions. The crawled data was �ltered for
safety, quality, and privacy.
Type of modality covered:  ☒ Text   ☒ Image   ☒ Video   □ Audio
□ Other If so, please specify..
Ask AI
Public Summary of Training Content – Midjourney https://docs.midjourney.com/hc/en-us/articles/4806708031...
4 of 7 08/09/2026, 17:08
Summary of the most relevant
domain names crawled:
The crawled data includes data from publicly
available online websites and sources that
encompass a large variety of content types and
languages. Toplevel domain names crawled
include: .com, .org, and .net as well as other
global sites.
Additional comments (optional):  N/A
Was data from user interactions
with the AI model (e.g. user input
and prompts) used to train the
model?
 ☒ Yes   □ No
Was data collected from user
interactions with the provider’s
other services or products used
to train the model?
 □ Yes   ☒ No
If yes, provide a general
description of the provider’s
services or products that were
used to collect the user data:
  N/A
Type of modality covered:  ☒ Text   ☒ Image   □ Video   □ Audio
□ Other If so, please specify...
Additional comments (optional):
Midjourney adheres to its Privacy Policy and
Terms of Service, as applicable, for user data
training.
Was synthetic AI-generated
data created by the provider or
on their behalf to train the
model?
 ☒ Yes   □ No
If yes, modality of the synthetic
data:   ☒ Text   ☒ Image   □ Video   □ Audio
□ Other If so, please specify...
Ask AI
Public Summary of Training Content – Midjourney https://docs.midjourney.com/hc/en-us/articles/4806708031...
5 of 7 08/09/2026, 17:08
If yes, specify the general-
purpose AI model(s) used to
generate the synthetic data if
available on the market:
Midjourney may use internal models to generate
synthetic data for training.
Information about other AI
models, including provider’s own
AI model(s) not available on the
market, used to generate
synthetic data to train the
model to which this Summary
applies:
  N/A
Additional comments (optional):  N/A
Have data sources other than
those described in Sections 2.1 to
2.5 been used to train the model?
 ☒ Yes   □ No
If yes, provide a narrative
description of these data sources
and the data:
  Midjourney used datasets that it has acquired
through its business operations.
Additional comments (optional):  N/A
Are you a Signatory to the Code
of Practice for general-purpose
AI models that includes
commitments to respect
reservations of rights from the
TDM exception or limitation?
 □ Yes   ☒ No
Ask AI
Public Summary of Training Content – Midjourney https://docs.midjourney.com/hc/en-us/articles/4806708031...
6 of 7 08/09/2026, 17:08
Describe the measures
implemented before model
training to respect reservations
of rights from the TDM exception
or limitation before and during
data collection, including the
opt-out protocols and solutions
honoured by the provider or, as
applicable, by third parties from
which datasets have been
obtained:

Midjourney implemented safety and quality
�ltering, content moderation, honored
robots.txt �les instructions where appropriate,
and conducted privacy processing to �lter or
remove sensitive personal information from
training data.
Additional comments (optional):  N/A
General description of measures
taken:
Midjourney adheres to applicable laws and
best practices with respect to removing illegal
content. Midjourney training data undergoes
safety �ltering to remove certain data with
known risk of containing child sexual abuse
material (CSAM) and other categories of
sensitive or disallowed content. See also rules
of conduct for its users at: https://
docs.midjourney.com/hc/en-us/
articles/32013696484109-Community-
Guidelines.
Other relevant information about
data processing (optional):   N/A
Midjourney Website Midjourney Discord Server
Ask AI
Public Summary of Training Content – Midjourney https://docs.midjourney.com/hc/en-us/articles/4806708031...
7 of 7 08/09/2026, 17:08