GPAI Ledger The public record of EU AI Act training-data summaries

GPAI LedgerGPAI Training Transparency tracker (AI Accountability Lab (AIAL)) › Capture 12 Sep 2026

ByteDance_Seedeam_5_0_2026_09_08 — capture 20260912T062129Z

Filed under AI Accountability Lab (AIAL) — GPAI Training Transparency tracker, the source this project captured it from; the document itself is the filing of the model named above.

Providerprovider not identified by this project
TargetAIAL archived copy — https://raw.githubusercontent.com/AIAccountabilityLab/gpai-training-transparency/c878ebbe5a3d7585857ed203f59f34551d7e72f7/public/archive/ByteDance_Seedeam_5_0_2026_09_08.pdf
Fetched (UTC)2026-09-12T06:21:29Z
Upstream commit8 Sep 2026 — c878ebbe5a3d (when this state began to stand in the upstream repository; this archive fetched it at the time above, not then)
Stored filebbfe07316d3bfeb78f67e43c8280e887102aa2cb4926db18e1a3c9fa25fa1e7e.pdf (251,620 bytes)
SHA-256bbfe07316d3bfeb78f67e43c8280e887102aa2cb4926db18e1a3c9fa25fa1e7e
OpenTimestamps proofbbfe07316d3bfeb78f67e43c8280e887102aa2cb4926db18e1a3c9fa25fa1e7e.pdf.20260912T062129Z.ots (calendar-attested; anchored in bitcoin over time)
Waybacknot saved
Prior capture of this target— first capture of this target

Verify: sha256sum bbfe07316d3bfeb78f67e43c8280e887102aa2cb4926db18e1a3c9fa25fa1e7e.pdf must equal the hash above (the filename IS the expected hash); ots verify bbfe07316d3bfeb78f67e43c8280e887102aa2cb4926db18e1a3c9fa25fa1e7e.pdf.20260912T062129Z.ots -f bbfe07316d3bfeb78f67e43c8280e887102aa2cb4926db18e1a3c9fa25fa1e7e.pdf (opentimestamps.org) proves the bytes existed no later than the attestation time — an upper bound on the capture time; the fetch time above is the archive's own record (a freshly captured proof reports 'pending' here: the calendars anchor within hours, but this archive only upgrades the stored proof to its anchor on a later run, so expect a day or two). ots verify needs a local Bitcoin Core node (a pruned one is fine); without one, ots info on the proof prints the attesting block height and merkle path to check on any block explorer.

Extracted text

Machine-extracted text (layout may be lost; the authoritative content is the stored file above).

PublicSummaryofTrainingDataContent
forSeedream5.0Pro
VersionoftheSummary: 1.0
Lastupdate: 31July2026
1. Generalinformation
1.1 Provideridentification
Providernameandcontact
details:
ByteDanceNexusAIPte.Ltd(ByteDance)

Authorisedrepresentative
nameandcontactdetails:
MikrosInformationTechnologyIrelandLimited
1.2 Modelidentification
Versionedmodelname(s): TheSeedream5.0Profamilyofmodels
Modeldependencies: TheSeedream5.0ProfamilyincludesSeeDream5.0ProandSeedream5.0
Lite.Bothmodelsarebasedonthesamearchitecture.
Dateofplacementofthe
modelontheUnion
market:
July2026
1.3 Modalities,overalltrainingdatasizeandothercharacteristics
Modality

Trainingdatasize

Typesofcontent

Text

Morethan10trilliontokens

Seedream5.0Pro'strainingdataset
comprisespairedimage-textdatasets
andinterweavedimage-textdata,
includingimageswithcorresponding
textualdescriptionsanddocuments
withinterweavedimagesandtext.
Image

Morethan1billionimages

Seedream5.0Pro'strainingdataset
comprisespairedimage-textdatasets
andinterweavedimage-textdata,
includingimageswithcorresponding
textualdescriptionsanddocuments
withinterweavedimagesandtext.
Audio N/A

N/A
Video N/A N/A
Other N/A N/A
Latestdateofdata
acquisition/collectionfor
modeltraining:
June2026

Descriptionofthelinguistic
characteristicsoftheoverall
trainingdata:
Nospecificgeographicregionwasintentionallyexcludedfromthe
datacollectionprocess.


Otherrelevantcharacteristics
oftheoveralltrainingdata:
Seedream5.0Pro'strainingdatasetcomprisespairedimage-text
datasetsandinterweavedimage-textdata,includingimageswith
correspondingtextualdescriptionsanddocumentswithinterweaved
imagesandtext.
Formodeltrainingpurposes,ByteDanceusesalarge-scalemixtureof
publiclyavailabledata,licenseddata,andsyntheticdata,designedto
ensurebroadcoverageofdomainsandcontexts.
Thedatasethasbeencuratedwiththeobjectiveofmaximizing
representativenessandrobustness,whileapplyingfiltering
mechanismstoremovelow-qualityorharmfulcontentwhere
appropriate.

Additionalcomments
(optional):
N/A
2. ListofDataSources
2.1 Publiclyavailabledatasets
Haveyouusedpubliclyavailable
datasetstotrainthemodel?

Yes
Ifyes,specifythemodality(ies)
ofthecontentcoveredbythe
datasetsconcerned:
Image

Listoflargepubliclyavailable
datasets:
Thetrainingdatacorpuscomprisesamixofpubliclyavailableand
licenseddata.
Thismayincludeimagedatasetsmadeavailablebythirdparties
throughpublicrepositoriesandonlineplatforms.Asdescribed
above,suchdatasetsarecuratedandpre-processedwiththe
objectiveofmaximizingrepresentativenessandrobustness,while
applyingfilteringmechanismstoremovelow-qualityorharmful
contentwhereappropriate.
Generaldescriptionofother
publiclyavailabledatasetsnot
listedabove:
Seethedescriptionoflinguisticcharacteristicsandotherrelevant
characteristicsofoveralltrainingdatainSection1.
Additionalcomments(optional): N/A
2.2 Privatenon-publiclyavailabledatasetsobtainedfromthird
parties
2.2.1 Datasetscommerciallylicensedbyrightsholdersortheir
representatives
Haveyouconcludedtransactionalcommerciallicensing
agreement(s)withrightsholder(s)orwiththeir
representatives?
Yes

Image

Ifyes,specifythemodality(ies)ofthecontentcoveredby
thedatasetsconcerned:
2.2.2 Privatedatasetsobtainedfromotherthirdparties
Haveyouobtainedprivatedatasetsfrom
thirdpartiesthatarenotlicensedas
describedinSection2.2.1,suchasdata
obtainedfromprovidersofprivatedatabases,
ordataintermediaries?
Yes
Ifyes,specifythemodality(ies)ofthecontent
coveredbythedatasetsconcerned:
Image
Ifpubliclyknown,listprivatedatasets
obtainedfromotherthirdparties:
N/A
Generaldescriptionofnon-publiclyknown
privatedatasetsobtainedfromthirdparties
Wehaveobtaineddatasetsfromthird-partiesona
licensedbasis.Suchdatasetsprimarilycomprise
contentacrosstextandimagemodalities.
Additionalcomments(optional): N/A
2.3 Datacrawledandscrapedfromonlinesources
Werecrawlersusedbytheprovideror
onbehalfof?
Yes
Ifyes,specifycrawler
name(s)/identifier(s):
Bytespider
Purposesofthecrawler(s): InthecontextofByteDance'sAImodeldevelopment,
Bytespiderisusedtocollectpubliclyaccessibleimage,video
andaudiodatatosupporttraining,testing,andvalidationof
theAImodel.
Generaldescriptionofcrawler
behaviour:
ByteDance'scrawlerisconfiguredtoavoidoverloading
websitesandtooperateinamannerconsistentwithethical
webscrapingpractices.Ourcrawlerisdesignedtorespect
robots.txtinstructions,andnottocircumventcontrol
measuressuchaspaywallsortoaccesspassword-protected
content.
Periodofdatacollection: UptoJune2026

Comprehensivedescriptionofthetype
ofcontentandonlinesourcescrawled:
Asabove,wemayuseourcrawlertocollectpublicly
accessibledata.Thecrawledcontentprimarilyconsistsof
videos,imagesandaudio.
Typeofmodalitycovered: Image
Summaryofthemostrelevantdomain
namescrawled:
Themostrelevantdomainscrawledincludesresources
spanningnaturalscenes,objects,peopleanddiagrams.
Additionalcomments(optional): N/A
2.4 Userdata
WasdatafromuserinteractionswiththeAImodel
(e.g.userinputandprompts)usedtotrainthe
model?
No
Wasdatacollectedfromuserinteractionswiththe
provider’sotherservicesorproductsusedtotrain
themodel?

No
Ifyes,provideageneraldescriptionofthe
provider’sservicesorproductsthatwereusedto
collecttheuserdata:
N/A
Typeofmodalitycovered: N/A
Additionalcomments(optional): N/A
2.5 Syntheticdata
WassyntheticAI-generateddatacreatedbythe
providerorontheirbehalftotrainthemodel?
Yes
Ifyes,modalityofthesyntheticdata: Textandimage
Ifyes,specifythegeneral-purposeAImodel(s)used
togeneratethesyntheticdataifavailableonthe
market:
Syntheticdatawasgeneratedusingarangeof
general-purposeAImodels,includingimage
generationmodels,visionlanguagemodelsand
largelanguagemodels.

InformationaboutotherAImodels,including
provider’sownAImodel(s)notavailableonthe
market,usedtogeneratesyntheticdatatotrainthe
modeltowhichthisSummaryapplies:
Wemayusefine-tunedversionsofinternal
modelstogeneratesyntheticdata.
Additionalcomments(optional): N/A
2.6 Othersourcesofdata
Havedatasourcesotherthanthosedescribedin
Sections2.1to2.5beenusedtotrainthemodel?
No
Ifyes,provideanarrativedescriptionofthesedata
sourcesandthedata:
N/A
Additionalcomments(optional): N/A
3. Dataprocessingaspects
3.1 Respectofreservationofrightsfromtextanddatamining
exceptionorlimitation
AreyouaSignatorytotheCodeofPracticeforgeneral-
purposeAImodelsthatincludescommitmentstorespect
reservationsofrightsfromtheTDMexceptionorlimitation?
No
Describethemeasuresimplementedbeforemodeltraining
torespectreservationsofrightsfromtheTDMexceptionor
limitationbeforeandduringdatacollection,includingthe
opt-outprotocolsandsolutionshonouredbytheprovider
or,asapplicable,bythirdpartiesfromwhichdatasetshave
beenobtained:

ByteDance'scrawlerisconfiguredtoavoid
overloadingwebsitesandtooperateina
mannerconsistentwithethicalweb
scrapingpractices.Ourcrawleris
designedtorespectrobots.txt
instructions,andnottocircumvent
controlmeasuressuchaspaywallsorto
accesspassword-protectedcontent.
Additionalcomments(optional): N/A
3.2 Removalofillegalcontent
Generaldescriptionofmeasurestaken:

Asmentioned,weappliedpreprocessingandfiltering
methods,includingfilteringmodels,toexcludeunsafeor
harmfulcontent.
3.3 Otherinformation(optional)
Otherrelevantinformationaboutdata
processing(optional):
N/A