GPAI Ledger The public record of EU AI Act training-data summaries

GPAI LedgerGPAI Training Transparency tracker (AI Accountability Lab (AIAL)) › Capture 12 Sep 2026

ByteDance_Seedance_2_0_2026_09_08 — capture 20260912T062122Z

Filed under AI Accountability Lab (AIAL) — GPAI Training Transparency tracker, the source this project captured it from; the document itself is the filing of the model named above.

Providerprovider not identified by this project
TargetAIAL archived copy — https://raw.githubusercontent.com/AIAccountabilityLab/gpai-training-transparency/c878ebbe5a3d7585857ed203f59f34551d7e72f7/public/archive/ByteDance_Seedance_2_0_2026_09_08.pdf
Fetched (UTC)2026-09-12T06:21:22Z
Upstream commit8 Sep 2026 — c878ebbe5a3d (when this state began to stand in the upstream repository; this archive fetched it at the time above, not then)
Stored fileae9aebb04101dab496b15addb920ff2725ade15ad96efbce5733b78274f1da07.pdf (246,255 bytes)
SHA-256ae9aebb04101dab496b15addb920ff2725ade15ad96efbce5733b78274f1da07
OpenTimestamps proofae9aebb04101dab496b15addb920ff2725ade15ad96efbce5733b78274f1da07.pdf.20260912T062122Z.ots (calendar-attested; anchored in bitcoin over time)
Waybacknot saved
Prior capture of this target— first capture of this target

Verify: sha256sum ae9aebb04101dab496b15addb920ff2725ade15ad96efbce5733b78274f1da07.pdf must equal the hash above (the filename IS the expected hash); ots verify ae9aebb04101dab496b15addb920ff2725ade15ad96efbce5733b78274f1da07.pdf.20260912T062122Z.ots -f ae9aebb04101dab496b15addb920ff2725ade15ad96efbce5733b78274f1da07.pdf (opentimestamps.org) proves the bytes existed no later than the attestation time — an upper bound on the capture time; the fetch time above is the archive's own record (a freshly captured proof reports 'pending' here: the calendars anchor within hours, but this archive only upgrades the stored proof to its anchor on a later run, so expect a day or two). ots verify needs a local Bitcoin Core node (a pruned one is fine); without one, ots info on the proof prints the attesting block height and merkle path to check on any block explorer.

Extracted text

Machine-extracted text (layout may be lost; the authoritative content is the stored file above).

PublicSummaryofTrainingDataContent
forSeedance2.0
VersionoftheSummary: 1.0
Lastupdate: 31July2026
1. Generalinformation
1.1 Provideridentification
Providernameand
contactdetails:
ByteDanceNexusAIPte.Ltd(ByteDance)

Authorised
representativenameand
contactdetails:
MikrosInformationTechnologyIrelandLimited
1.2 Modelidentification
Versionedmodel
name(s):
Seedance2.0
Modeldependencies: N/A
Dateofplacementofthe
modelontheUnion
market:
March2026
1.3 Modalities,overalltrainingdatasizeandothercharacteristics
Modality

Trainingdatasize

Typesofcontent
Text 1billionto10trilliontokens

 Themodelwaspre-trainedusingalarge-
scale,diversecollectionofdataspanning
awide-rangeofdomainsandmodalities,
includingtextualcaptions.
Image Morethan1billionimages Themodelwaspre-trainedusingalarge-
scale,diversecollectionofdataspanning
awide-rangeofdomainsandmodalities,
includinginfographicsandimagery.
Audio Morethan1millionhours Mostaudiocontentoriginatesfromthe
associatedvideofiles,andonlyalimited
subsetconsistsofstandaloneaudio.
Video Morethan1millionhours Seedance2.0'strainingdataset
comprisedvideowithaccordinglytextual
description.
Other N/A N/A
Latestdateofdata
acquisition/collectionformodel
training:
February2026

Descriptionofthelinguistic
characteristicsoftheoverall
trainingdata:
Nospecificgeographicregionwasintentionallyexcludedfromthe
datacollectionprocess.Videodatacontainsadiverserangeof
globallanguages,includinglanguagesfromwithinEurope.

Otherrelevantcharacteristicsof
theoveralltrainingdata:
Seedance2.0'strainingdatasetcomprisedclipsofvideowith
accordinglytextualdescription.
Formodeltrainingpurposes,ByteDanceusesalarge-scalemixture
ofpubliclyavailabledata,licenseddata,andsyntheticdata,
designedtoensurebroadcoverageofdomainsandcontexts.
Thedatasethasbeencuratedwiththeobjectiveofmaximizing
representativenessandrobustness,whileapplyingfiltering
mechanismstoremovelow-qualityorharmfulcontentwhere
appropriate.
Additionalcomments(optional): N/A
2. ListofDataSources

2.1 Publiclyavailabledatasets
Haveyouusedpubliclyavailable
datasetstotrainthemodel?
Yes
Ifyes,specifythemodality(ies)
ofthecontentcoveredbythe
datasetsconcerned:
Image,video,audio

Listoflargepubliclyavailable
datasets:
Thetrainingdatacorpuscomprisesamixofpubliclyavailableand
licenseddata.
Thismayincludeimage,videoandaudiodatasetsmadeavailableby
thirdpartiesthroughpublicrepositoriesandonlineplatforms.As
describedabove,suchdatasetsarecuratedandpre-processedwith
theobjectiveofmaximizingrepresentativenessandrobustness,
whileapplyingfilteringmechanismstoremovelow-qualityor
harmfulcontentwhereappropriate.
Generaldescriptionofother
publiclyavailabledatasetsnot
listedabove:
Seethedescriptionoflinguisticcharacteristicsandotherrelevant
characteristicsofoveralltrainingdatainSection1.
Additionalcomments(optional): N/A
2.2 Privatenon-publiclyavailabledatasetsobtainedfromthird
parties
2.2.1 Datasetscommerciallylicensedbyrightsholdersortheir
representatives
Haveyouconcludedtransactional
commerciallicensingagreement(s)
withrightsholder(s)orwiththeir
representatives?
Yes

Ifyes,specifythemodality(ies)ofthe
contentcoveredbythedatasets
concerned:
Image,videoandaudio
2.2.2 Privatedatasetsobtainedfromotherthirdparties

Haveyouobtainedprivatedatasets
fromthirdpartiesthatarenot
licensedasdescribedinSection2.2.1,
suchasdataobtainedfromproviders
ofprivatedatabases,ordata
intermediaries?
Yes
Ifyes,specifythemodality(ies)ofthe
contentcoveredbythedatasets
concerned:
Image,videoandaudio
Ifpubliclyknown,listprivatedatasets
obtainedfromotherthirdparties:
N/A
Generaldescriptionofnon-publicly
knownprivatedatasetsobtainedfrom
thirdparties
Wehaveobtaineddatasetsfromthird-partiesonalicensed
basis.Suchdatasetscomprisecontentacrossimage,videoand
audiomodalities.
Additionalcomments(optional): N/A
2.3 Datacrawledandscrapedfromonlinesources
Werecrawlersusedbytheprovideror
onbehalfof?
Yes
Ifyes,specifycrawler
name(s)/identifier(s):
Bytespider
Purposesofthecrawler(s): InthecontextofByteDance'sAImodeldevelopment,
Bytespiderisusedtocollectpubliclyaccessibleimage,video
andaudiodatatosupporttraining,testing,andvalidationof
theAImodel.
Generaldescriptionofcrawler
behaviour:

ByteDance'scrawlerisconfiguredtoavoidoverloading
websitesandtooperateinamannerconsistentwithethical
webscrapingpractices.Ourcrawlerisdesignedtorespect
robots.txtinstructions,andnottocircumventcontrolmeasures
suchaspaywallsortoaccesspassword-protectedcontent.
Periodofdatacollection: 2025
Comprehensivedescriptionofthe
typeofcontentandonlinesources
crawled:
Asabove,wemayuseourcrawlertocollectpubliclyaccessible
data.Thecrawledcontentprimarilyconsistsofimagesand
video.
Typeofmodalitycovered: Imageandvideo

Summaryofthemostrelevant
domainnamescrawled:
Themostrelevantdomainscrawledincludesresources
spanningnaturalscenes,objects,peopleanddiagrams.
Additionalcomments(optional): N/A
2.4 Userdata
WasdatafromuserinteractionswiththeAI
model(e.g.userinputandprompts)usedtotrain
themodel?
No
Wasdatacollectedfromuserinteractionswith
theprovider’sotherservicesorproductsused
totrainthemodel?

No
Ifyes,provideageneraldescriptionofthe
provider’sservicesorproductsthatwereused
tocollecttheuserdata:
N/A
Typeofmodalitycovered: N/A
Additionalcomments(optional): N/A
2.5 Syntheticdata
WassyntheticAI-generateddatacreatedbythe
providerorontheirbehalftotrainthemodel?
Yes
Ifyes,modalityofthesyntheticdata: Text(descriptionofvideosandimages)
Ifyes,specifythegeneral-purposeAImodel(s)
usedtogeneratethesyntheticdataifavailable
onthemarket:
Syntheticdatawasgeneratedusingarangeof
general-purposeAImodels,includingvision
languagemodelsandlargelanguagemodels.
InformationaboutotherAImodels,including
provider’sownAImodel(s)notavailableonthe
market,usedtogeneratesyntheticdatatotrain
themodeltowhichthisSummaryapplies:
Wemayusefine-tunedversionsofinternalmodels
togeneratesyntheticdata.
Additionalcomments(optional): N/A
2.6 Othersourcesofdata

Havedatasourcesotherthanthosedescribedin
Sections2.1to2.5beenusedtotrainthemodel?
No
Ifyes,provideanarrativedescriptionofthese
datasourcesandthedata:
N/A
Additionalcomments(optional): N/A
3. Dataprocessingaspects
3.1 Respectofreservationofrightsfromtextanddatamining
exceptionorlimitation
AreyouaSignatorytotheCodeofPracticeforgeneral-
purposeAImodelsthatincludescommitmentsto
respectreservationsofrightsfromtheTDMexception
orlimitation?
No
Describethemeasuresimplementedbeforemodel
trainingtorespectreservationsofrightsfromtheTDM
exceptionorlimitationbeforeandduringdata
collection,includingtheopt-outprotocolsand
solutionshonouredbytheprovideror,asapplicable,
bythirdpartiesfromwhichdatasetshavebeen
obtained:
ByteDance'scrawlerisconfiguredtoavoid
overloadingwebsitesandtooperateina
mannerconsistentwithethicalwebscraping
practices.Ourcrawlerisdesignedtorespect
robots.txtinstructions,andnottocircumvent
controlmeasuressuchaspaywallsorto
accesspassword-protectedcontent.
Additionalcomments(optional): N/A
3.2 Removalofillegalcontent
Generaldescriptionofmeasures
taken:
Asmentioned,weappliedpreprocessingandfilteringmethods,
includingfilteringmodels,toexcludeunsafeorharmfulcontent.
3.3 Otherinformation(optional)
Otherrelevantinformationabout
dataprocessing(optional):
N/A