GPAI Ledger The public record of EU AI Act training-data summaries

GPAI LedgerGPAI Training Transparency tracker (AI Accountability Lab (AIAL)) › Capture 12 Sep 2026

ByteDance_Seedance_2_5_2026_09_08 — capture 20260912T062125Z

Filed under AI Accountability Lab (AIAL) — GPAI Training Transparency tracker, the source this project captured it from; the document itself is the filing of the model named above.

Providerprovider not identified by this project
TargetAIAL archived copy — https://raw.githubusercontent.com/AIAccountabilityLab/gpai-training-transparency/c878ebbe5a3d7585857ed203f59f34551d7e72f7/public/archive/ByteDance_Seedance_2_5_2026_09_08.pdf
Fetched (UTC)2026-09-12T06:21:25Z
Upstream commit8 Sep 2026 — c878ebbe5a3d (when this state began to stand in the upstream repository; this archive fetched it at the time above, not then)
Stored file629ed017d4743e5f7a4959cff420a79536de7c1386db0e0852ea9ce0a19f54fc.pdf (248,271 bytes)
SHA-256629ed017d4743e5f7a4959cff420a79536de7c1386db0e0852ea9ce0a19f54fc
OpenTimestamps proof629ed017d4743e5f7a4959cff420a79536de7c1386db0e0852ea9ce0a19f54fc.pdf.20260912T062125Z.ots (calendar-attested; anchored in bitcoin over time)
Waybacknot saved
Prior capture of this target— first capture of this target

Verify: sha256sum 629ed017d4743e5f7a4959cff420a79536de7c1386db0e0852ea9ce0a19f54fc.pdf must equal the hash above (the filename IS the expected hash); ots verify 629ed017d4743e5f7a4959cff420a79536de7c1386db0e0852ea9ce0a19f54fc.pdf.20260912T062125Z.ots -f 629ed017d4743e5f7a4959cff420a79536de7c1386db0e0852ea9ce0a19f54fc.pdf (opentimestamps.org) proves the bytes existed no later than the attestation time — an upper bound on the capture time; the fetch time above is the archive's own record (a freshly captured proof reports 'pending' here: the calendars anchor within hours, but this archive only upgrades the stored proof to its anchor on a later run, so expect a day or two). ots verify needs a local Bitcoin Core node (a pruned one is fine); without one, ots info on the proof prints the attesting block height and merkle path to check on any block explorer.

Extracted text

Machine-extracted text (layout may be lost; the authoritative content is the stored file above).

PublicSummaryofTrainingDataContent
forSeedance2.5
VersionoftheSummary: 1.0
Lastupdate: 31July2026
1. Generalinformation
1.1 Provideridentification
Providernameandcontact
details:
ByteDanceNexusAIPte.Ltd(ByteDance)
Authorisedrepresentative
nameandcontactdetails:
MikrosInformationTechnologyIrelandLimited

1.2 Modelidentification
Versionedmodelname(s): Seedance2.5
Modeldependencies: N/A
Dateofplacementofthe
modelontheUnion
market:
July2026
1.3 Modalities,overalltrainingdatasizeandothercharacteristics
Modality

Trainingdatasize

Typesofcontent
Text

1billionto10trilliontokens Themodelwaspre-trainedusinga
large-scale,diversecollectionofdata
spanningawide-rangeofdomains

andmodalities,includingtextual
captions.
Image Morethan1billionimages Themodelwaspre-trainedusinga
large-scale,diversecollectionofdata
spanningawide-rangeofdomains
andmodalities,including
infographicsandimagery.
Audio Morethan1millionhours Mostaudiocontentoriginatesfrom
theassociatedvideofiles,andonlya
limitedsubsetconsistsofstandalone
audio.
Video Morethan1millionhours Seedance2.5'strainingdataset
comprisedvideoclips.
Other N/A N/A
Latestdateofdata
acquisition/collectionfor
modeltraining:
June2026

Descriptionofthelinguistic
characteristicsoftheoverall
trainingdata:
Nospecificgeographicregionwasintentionallyexcludedfromthe
datacollectionprocess.Videodatacontainsadiverserangeofglobal
languages,includinglanguagesfromwithinEurope.

Otherrelevantcharacteristics
oftheoveralltrainingdata:
Seedance2.5'strainingdatasetcomprisedvideoclips.
Formodeltrainingpurposes,ByteDanceusesalarge-scalemixtureof
publiclyavailabledata,licenseddata,andsyntheticdata,designedto
ensurebroadcoverageofdomainsandcontexts.
Thedatasethasbeencuratedwiththeobjectiveofmaximizing
representativenessandrobustness,whileapplyingfiltering
mechanismstoremovelow-qualityorharmfulcontentwhere
appropriate.
Additionalcomments
(optional):
N/A
2. ListofDataSources
2.1 Publiclyavailabledatasets

Haveyouusedpubliclyavailable
datasetstotrainthemodel?
Yes

Ifyes,specifythemodality(ies)
ofthecontentcoveredbythe
datasetsconcerned:
Imageandvideo

Listoflargepubliclyavailable
datasets:
Thetrainingdatacorpuscomprisesamixofpubliclyavailableand
licenseddata.
Thismayincludeimageandvideodatasetsmadeavailablebythird
partiesthroughpublicrepositoriesandonlineplatforms.As
describedabove,suchdatasetsarecuratedandpre-processedwith
theobjectiveofmaximizingrepresentativenessandrobustness,
whileapplyingfilteringmechanismstoremovelow-qualityor
harmfulcontentwhereappropriate.
Generaldescriptionofother
publiclyavailabledatasetsnot
listedabove:
Seethedescriptionoflinguisticcharacteristicsandotherrelevant
characteristicsofoveralltrainingdatainSection1.
Additionalcomments(optional): N/A
2.2 Privatenon-publiclyavailabledatasetsobtainedfromthird
parties
2.2.1 Datasetscommerciallylicensedbyrightsholdersortheir
representatives
Haveyouconcludedtransactional
commerciallicensingagreement(s)with
rightsholder(s)orwiththeir
representatives?
Yes

Ifyes,specifythemodality(ies)ofthe
contentcoveredbythedatasets
concerned:
Image,videoandaudio
2.2.2 Privatedatasetsobtainedfromotherthirdparties
Haveyouobtainedprivatedatasetsfrom
thirdpartiesthatarenotlicensedas
describedinSection2.2.1,suchasdata
Yes

obtainedfromprovidersofprivate
databases,ordataintermediaries?
Ifyes,specifythemodality(ies)ofthe
contentcoveredbythedatasets
concerned:
Image,videoandaudio
Ifpubliclyknown,listprivatedatasets
obtainedfromotherthirdparties:
N/A
Generaldescriptionofnon-publicly
knownprivatedatasetsobtainedfrom
thirdparties
Wehaveobtaineddatasetsfromthird-partiesonalicensed
basis.Suchdatasetscomprisecontentacrossimage,video
andaudiomodalities.
Additionalcomments(optional): N/A
2.3 Datacrawledandscrapedfromonlinesources
Werecrawlersusedbytheprovideror
onbehalfof?
Yes
Ifyes,specifycrawler
name(s)/identifier(s):
Bytespider
Purposesofthecrawler(s): InthecontextofByteDance'sAImodeldevelopment,
Bytespiderisusedtocollectpubliclyaccessibleimage,video
andaudiodatatosupporttraining,testing,andvalidationof
theAImodel.
Generaldescriptionofcrawler
behaviour:
ByteDance'scrawlerisconfiguredtoavoidoverloading
websitesandtooperateinamannerconsistentwithethical
webscrapingpractices.Ourcrawlerisdesignedtorespect
robots.txtinstructions,andnottocircumventcontrol
measuressuchaspaywallsortoaccesspassword-protected
content.
Periodofdatacollection: UptoJune2026
Comprehensivedescriptionofthetype
ofcontentandonlinesourcescrawled:
Asabove,wemayuseourcrawlertocollectpublicly
accessibledata.Thecrawledcontentprimarilyconsistsof
videos,imagesandaudio.
Typeofmodalitycovered: Video,imagesandaudio.
Summaryofthemostrelevantdomain
namescrawled:
Themostrelevantdomainscrawledincludesresources
spanningnaturalscenes,objects,peopleanddiagrams.

Additionalcomments(optional): N/A
2.4 Userdata
WasdatafromuserinteractionswiththeAImodel
(e.g.userinputandprompts)usedtotrainthe
model?
No
Wasdatacollectedfromuserinteractionswiththe
provider’sotherservicesorproductsusedtotrain
themodel?
No
Ifyes,provideageneraldescriptionofthe
provider’sservicesorproductsthatwereusedto
collecttheuserdata:
N/A
Typeofmodalitycovered: N/A
Additionalcomments(optional): N/A
2.5 Syntheticdata
WassyntheticAI-generateddatacreatedbythe
providerorontheirbehalftotrainthemodel?
Yes
Ifyes,modalityofthesyntheticdata: Text(descriptionofvideosandimages)
Ifyes,specifythegeneral-purposeAImodel(s)used
togeneratethesyntheticdataifavailableonthe
market:
Syntheticdatawasgeneratedusingarangeof
general-purposeAImodels,includingvision
languagemodelsandlargelanguagemodels.

InformationaboutotherAImodels,including
provider’sownAImodel(s)notavailableonthe
market,usedtogeneratesyntheticdatatotrainthe
modeltowhichthisSummaryapplies:
Wemayusefine-tunedversionsofinternal
modelstogeneratesyntheticdata.
Additionalcomments(optional): N/A
2.6 Othersourcesofdata
No

Havedatasourcesotherthanthosedescribedin
Sections2.1to2.5beenusedtotrainthemodel?
Ifyes,provideanarrativedescriptionofthesedata
sourcesandthedata:
N/A
Additionalcomments(optional): N/A
3. Dataprocessingaspects
3.1 Respectofreservationofrightsfromtextanddatamining
exceptionorlimitation
AreyouaSignatorytotheCodeofPracticeforgeneral-
purposeAImodelsthatincludescommitmentstorespect
reservationsofrightsfromtheTDMexceptionorlimitation?
No

Describethemeasuresimplementedbeforemodeltraining
torespectreservationsofrightsfromtheTDMexceptionor
limitationbeforeandduringdatacollection,includingthe
opt-outprotocolsandsolutionshonouredbytheprovider
or,asapplicable,bythirdpartiesfromwhichdatasetshave
beenobtained:

ByteDance'scrawlerisconfiguredtoavoid
overloadingwebsitesandtooperateina
mannerconsistentwithethicalweb
scrapingpractices.Ourcrawleris
designedtorespectrobots.txt
instructions,andnottocircumvent
controlmeasuressuchaspaywallsorto
accesspassword-protectedcontent.
Additionalcomments(optional): N/A
3.2 Removalofillegalcontent
Generaldescriptionofmeasurestaken: Asmentioned,weappliedpreprocessingandfiltering
methods,includingfilteringmodels,toexcludeunsafeor
harmfulcontent.
3.3 Otherinformation(optional)
Otherrelevantinformationaboutdata
processing(optional):
N/A