GPAI Ledger The public record of EU AI Act training-data summaries

GPAI LedgerGPAI Training Transparency tracker (AI Accountability Lab (AIAL)) › Capture 12 Sep 2026

ByteDance_Seed_2_0_2026_09_08 — capture 20260912T062118Z

Filed under AI Accountability Lab (AIAL) — GPAI Training Transparency tracker, the source this project captured it from; the document itself is the filing of the model named above.

Providerprovider not identified by this project
TargetAIAL archived copy — https://raw.githubusercontent.com/AIAccountabilityLab/gpai-training-transparency/c878ebbe5a3d7585857ed203f59f34551d7e72f7/public/archive/ByteDance_Seed_2_0_2026_09_08.pdf
Fetched (UTC)2026-09-12T06:21:17Z
Upstream commit8 Sep 2026 — c878ebbe5a3d (when this state began to stand in the upstream repository; this archive fetched it at the time above, not then)
Stored file0b6178953deaed3aec71895c95bc62d30cd97df6368327115ec8d1506e33e3ce.pdf (251,504 bytes)
SHA-2560b6178953deaed3aec71895c95bc62d30cd97df6368327115ec8d1506e33e3ce
OpenTimestamps proof0b6178953deaed3aec71895c95bc62d30cd97df6368327115ec8d1506e33e3ce.pdf.20260912T062118Z.ots (calendar-attested; anchored in bitcoin over time)
Waybacknot saved
Prior capture of this target— first capture of this target

Verify: sha256sum 0b6178953deaed3aec71895c95bc62d30cd97df6368327115ec8d1506e33e3ce.pdf must equal the hash above (the filename IS the expected hash); ots verify 0b6178953deaed3aec71895c95bc62d30cd97df6368327115ec8d1506e33e3ce.pdf.20260912T062118Z.ots -f 0b6178953deaed3aec71895c95bc62d30cd97df6368327115ec8d1506e33e3ce.pdf (opentimestamps.org) proves the bytes existed no later than the attestation time — an upper bound on the capture time; the fetch time above is the archive's own record (a freshly captured proof reports 'pending' here: the calendars anchor within hours, but this archive only upgrades the stored proof to its anchor on a later run, so expect a day or two). ots verify needs a local Bitcoin Core node (a pruned one is fine); without one, ots info on the proof prints the attesting block height and merkle path to check on any block explorer.

Extracted text

Machine-extracted text (layout may be lost; the authoritative content is the stored file above).

PublicSummaryofTrainingDataContent
forSeed2.0Pro
VersionoftheSummary: 1.0
Lastupdate: 31July2026
1. Generalinformation
1.1 Provideridentification
Providernameandcontact
details:
ByteDanceNexusAIPte.Ltd(ByteDance)
Authorisedrepresentative
nameandcontactdetails:
MikrosInformationTechnologyIrelandLimited
1.2 Modelidentification
Versionedmodelname(s): Seed2.0-Pro(ModelnameonByteplus:Dola-Seed-2.0-Pro)
Modeldependencies: N/A
Dateofplacementofthe
modelontheUnion
market:
March2026
1.3 Modalities,overalltrainingdatasizeandothercharacteristics
Modality

Trainingdatasize

Typesofcontent
Text

Morethan10trilliontokens Themodelwaspre-trainedusinga
large-scale,diversecollectionofdata
spanningawide-rangeofdomains
andmodalities,includingdatasets

containinginformationrepresented
aswrittenlanguage.
Image Morethan1billionimages

Themodelwaspre-trainedusinga
large-scale,diversecollectionofdata
spanningawide-rangeofdomains
andmodalities,includingdatasets
containingstaticandvisual
information(suchasphotographsand
diagrams).
Audio N/A

N/A

Video 10,000to1millionhours Themodelwaspre-trainedusinga
large-scale,diversecollectionofdata
spanningawide-rangeofdomains
andmodalities,includingdatasets
containingvideoclips.
Other N/A N/A
Latestdateofdata
acquisition/collectionfor
modeltraining:
February2026

Descriptionofthelinguistic
characteristicsoftheoverall
trainingdata:
Themodel'strainingdatacontainsadiverserangeofgloballanguages,
includinglanguagesfromwithinEurope.

Otherrelevantcharacteristics
oftheoveralltrainingdata:
Themodelwaspre-trainedusingalarge-scale,diversecollectionof
dataspanningawide-rangeofdomainsandmodalities,including
datasetscontaininginformationrepresentedaswrittenlanguage,static
andvisualinformation(suchasphotographsanddiagrams)andvideo
clips.
Seed2.0Pro'strainingdatasetwascarefullycuratedtomaximize
diversity,representativeness,androbustness,whileappropriate
filteringmechanismsareappliedtofilterforqualityandsafetyin
relationtolow-qualityandharmfulcontent.
Additionalcomments
(optional):
N/A
2. ListofDataSources

2.1 Publiclyavailabledatasets
Haveyouusedpubliclyavailable
datasetstotrainthemodel?
Yes

Ifyes,specifythemodality(ies)
ofthecontentcoveredbythe
datasetsconcerned:
Text,imageandvideo
Listoflargepubliclyavailable
datasets:
Wehaveusedpubliclyavailabledatasetsacrossvarioussectorsand
stagesofthemodeldevelopmentpipeline.Thesepubliclyavailable
datasetsaresubjecttopre-processingsuchasdatacleaning,
deduplication,filtering,selection,etctoexcludeunsafeorharmful
content.
Generaldescriptionofother
publiclyavailabledatasetsnot
listedabove:
Seethedescriptionoflinguisticcharacteristicsandotherrelevant
characteristicsofoveralltrainingdatainSection1.
Additionalcomments(optional): N/A
2.2 Privatenon-publiclyavailabledatasetsobtainedfromthird
parties
2.2.1 Datasetscommerciallylicensedbyrightsholdersortheir
representatives
Haveyouconcludedtransactionalcommerciallicensing
agreement(s)withrightsholder(s)orwiththeir
representatives?
Yes

Ifyes,specifythemodality(ies)ofthecontentcoveredby
thedatasetsconcerned:
Textandimage
2.2.2 Privatedatasetsobtainedfromotherthirdparties
Haveyouobtainedprivatedatasetsfrom
thirdpartiesthatarenotlicensedas
describedinSection2.2.1,suchasdata
obtainedfromprovidersofprivatedatabases,
ordataintermediaries?
Yes

Ifyes,specifythemodality(ies)ofthecontent
coveredbythedatasetsconcerned:
Textandimage
Ifpubliclyknown,listprivatedatasets
obtainedfromotherthirdparties:
N/A

Generaldescriptionofnon-publiclyknown
privatedatasetsobtainedfromthirdparties
Wehaveobtaineddatasetsfromthird-partiesona
licensedbasis.Suchdatasetsprimarilycomprise
contentacrosstextandimagemodalities.
Additionalcomments(optional): N/A
2.3 Datacrawledandscrapedfromonlinesources
Werecrawlersusedbytheprovideror
onbehalfof?
Yes
Ifyes,specifycrawler
name(s)/identifier(s):
Bytespider
Purposesofthecrawler(s): InthecontextofByteDance'sAImodeldevelopment,
Bytespiderisusedtocollectpubliclyaccessibledata,
includingtextandimagedatatosupporttraining,testing,
andvalidationoftheAImodel.
Generaldescriptionofcrawler
behaviour:
ByteDance'scrawlerisconfiguredtoavoidoverloading
websitesandtooperateinamannerconsistentwithethical
webscrapingpractices.Ourcrawlerisdesignedtorespect
robots.txtinstructions,andnottocircumventcontrol
measuressuchaspaywallsortoaccesspassword-protected
content.
Periodofdatacollection: ApproximatelyJanuary2022toJune2024.
Comprehensivedescriptionofthetype
ofcontentandonlinesourcescrawled:
Asabove,wemayuseourcrawlertocollectpublicly
accessibledata.Thecrawledcontentprimarilyconsistsof
textandimageandincludesavarietyofcontenttypes
includingacademic,mathematical,code-relatedandgeneral-
purposecontent.
Typeofmodalitycovered: Textandimage
Summaryofthemostrelevantdomain
namescrawled:
Crawledcontentincludesavarietyofpubliclyaccessible
onlinematerialsacrossvarioussectors,suchaseducation,
government,technologyandresearchsectors.

Additionalcomments(optional): N/A
2.4 Userdata
WasdatafromuserinteractionswiththeAImodel
(e.g.userinputandprompts)usedtotrainthe
model?
No
Wasdatacollectedfromuserinteractionswiththe
provider’sotherservicesorproductsusedtotrain
themodel?

Yes
Ifyes,provideageneraldescriptionofthe
provider’sservicesorproductsthatwereusedto
collecttheuserdata:
ByteDance'smodelsaretrainedusinga
proprietarymixofdatasets,whichtendtobe
large-scaleanddiverse.Suchindustry-standard
datasetstypicallyincludeamixtureofpublicly
availabledocumentsasdescribedinthis
disclosure.Ourdatasetsmayalsoincludedata
obtainedinaccordancewiththerelevantterms
andconditions,privacypolicyandpursuantto
relevantuser-controlsasapplicable.
Typeofmodalitycovered: Seeabove
Additionalcomments(optional): Weusedatafilteringprocessestoreduce
personalinformationfromtrainingdata,andto
reducetheamountofpersonaldatainour
trainingdata.
2.5 Syntheticdata
WassyntheticAI-generateddatacreatedbythe
providerorontheirbehalftotrainthemodel?
Yes
Ifyes,modalityofthesyntheticdata: Textandimage
Ifyes,specifythegeneral-purposeAImodel(s)used
togeneratethesyntheticdataifavailableonthe
market:
Syntheticdatawasgeneratedusingarangeof
general-purposeAImodels,includingvision
languagemodels,imagegenerationmodelsand
largelanguagemodels.

InformationaboutotherAImodels,including
provider’sownAImodel(s)notavailableonthe
market,usedtogeneratesyntheticdatatotrainthe
modeltowhichthisSummaryapplies:
Wemayusefine-tunedversionsofinternal
modelstogeneratesyntheticdata.

Additionalcomments(optional): N/A
2.6 Othersourcesofdata
Havedatasourcesotherthanthosedescribedin
Sections2.1to2.5beenusedtotrainthemodel?
Yes
Ifyes,provideanarrativedescriptionofthesedata
sourcesandthedata:
Weworkedwithourvendorstocreatelabelled
datasetssoastoimprovethemodel'scapability
oncertaintasks,suchasreasoning,codingand
knowledgetasks.
Additionalcomments(optional): N/A
3. Dataprocessingaspects
3.1 Respectofreservationofrightsfromtextanddatamining
exceptionorlimitation
AreyouaSignatorytotheCodeofPracticeforgeneral-
purposeAImodelsthatincludescommitmentstorespect
reservationsofrightsfromtheTDMexceptionorlimitation?
No

Describethemeasuresimplementedbeforemodeltraining
torespectreservationsofrightsfromtheTDMexceptionor
limitationbeforeandduringdatacollection,includingthe
opt-outprotocolsandsolutionshonouredbytheprovider
or,asapplicable,bythirdpartiesfromwhichdatasetshave
beenobtained:
ByteDance'scrawlerisconfiguredtoavoid
overloadingwebsitesandtooperateina
mannerconsistentwithethicalweb
scrapingpractices.Ourcrawleris
designedtorespectrobots.txt
instructions,andnottocircumvent
controlmeasuressuchaspaywallsorto
accesspassword-protectedcontent.
Additionalcomments(optional): N/A
3.2 Removalofillegalcontent

Generaldescriptionofmeasurestaken: Asmentioned,weappliedpreprocessingandfiltering
methods,includingfilteringmodelsfocusedonunsafeor
harmfulcontent.
3.3 Otherinformation(optional)
Otherrelevantinformationaboutdata
processing(optional):
N/A