GPAI Ledger › GPAI Training Transparency tracker (AI Accountability Lab (AIAL)) › Capture 12 Sep 2026
ByteDance_Seed_2_0_2026_09_08 — capture 20260912T062118Z
Filed under AI Accountability Lab (AIAL) — GPAI Training Transparency tracker, the source this project captured it from; the document itself is the filing of the model named above.
| Provider | provider not identified by this project |
|---|---|
| Target | AIAL archived copy — https://raw.githubusercontent.com/AIAccountabilityLab/gpai-training-transparency/c878ebbe5a3d7585857ed203f59f34551d7e72f7/public/archive/ByteDance_Seed_2_0_2026_09_08.pdf |
| Fetched (UTC) | 2026-09-12T06:21:17Z |
| Upstream commit | 8 Sep 2026 — c878ebbe5a3d (when this state began to stand in the upstream repository; this archive fetched it at the time above, not then) |
| Stored file | 0b6178953deaed3aec71895c95bc62d30cd97df6368327115ec8d1506e33e3ce.pdf (251,504 bytes) |
| SHA-256 | 0b6178953deaed3aec71895c95bc62d30cd97df6368327115ec8d1506e33e3ce |
| OpenTimestamps proof | 0b6178953deaed3aec71895c95bc62d30cd97df6368327115ec8d1506e33e3ce.pdf.20260912T062118Z.ots (calendar-attested; anchored in bitcoin over time) |
| Wayback | not saved |
| Prior capture of this target | — first capture of this target |
Verify: sha256sum 0b6178953deaed3aec71895c95bc62d30cd97df6368327115ec8d1506e33e3ce.pdf must equal the hash above (the filename IS the expected hash); ots verify 0b6178953deaed3aec71895c95bc62d30cd97df6368327115ec8d1506e33e3ce.pdf.20260912T062118Z.ots -f 0b6178953deaed3aec71895c95bc62d30cd97df6368327115ec8d1506e33e3ce.pdf (opentimestamps.org) proves the bytes existed no later than the attestation time — an upper bound on the capture time; the fetch time above is the archive's own record (a freshly captured proof reports 'pending' here: the calendars anchor within hours, but this archive only upgrades the stored proof to its anchor on a later run, so expect a day or two). ots verify needs a local Bitcoin Core node (a pruned one is fine); without one, ots info on the proof prints the attesting block height and merkle path to check on any block explorer.
Extracted text
Machine-extracted text (layout may be lost; the authoritative content is the stored file above).
PublicSummaryofTrainingDataContent forSeed2.0Pro VersionoftheSummary: 1.0 Lastupdate: 31July2026 1. Generalinformation 1.1 Provideridentification Providernameandcontact details: ByteDanceNexusAIPte.Ltd(ByteDance) Authorisedrepresentative nameandcontactdetails: MikrosInformationTechnologyIrelandLimited 1.2 Modelidentification Versionedmodelname(s): Seed2.0-Pro(ModelnameonByteplus:Dola-Seed-2.0-Pro) Modeldependencies: N/A Dateofplacementofthe modelontheUnion market: March2026 1.3 Modalities,overalltrainingdatasizeandothercharacteristics Modality Trainingdatasize Typesofcontent Text Morethan10trilliontokens Themodelwaspre-trainedusinga large-scale,diversecollectionofdata spanningawide-rangeofdomains andmodalities,includingdatasets containinginformationrepresented aswrittenlanguage. Image Morethan1billionimages Themodelwaspre-trainedusinga large-scale,diversecollectionofdata spanningawide-rangeofdomains andmodalities,includingdatasets containingstaticandvisual information(suchasphotographsand diagrams). Audio N/A N/A Video 10,000to1millionhours Themodelwaspre-trainedusinga large-scale,diversecollectionofdata spanningawide-rangeofdomains andmodalities,includingdatasets containingvideoclips. Other N/A N/A Latestdateofdata acquisition/collectionfor modeltraining: February2026 Descriptionofthelinguistic characteristicsoftheoverall trainingdata: Themodel'strainingdatacontainsadiverserangeofgloballanguages, includinglanguagesfromwithinEurope. Otherrelevantcharacteristics oftheoveralltrainingdata: Themodelwaspre-trainedusingalarge-scale,diversecollectionof dataspanningawide-rangeofdomainsandmodalities,including datasetscontaininginformationrepresentedaswrittenlanguage,static andvisualinformation(suchasphotographsanddiagrams)andvideo clips. Seed2.0Pro'strainingdatasetwascarefullycuratedtomaximize diversity,representativeness,androbustness,whileappropriate filteringmechanismsareappliedtofilterforqualityandsafetyin relationtolow-qualityandharmfulcontent. Additionalcomments (optional): N/A 2. ListofDataSources 2.1 Publiclyavailabledatasets Haveyouusedpubliclyavailable datasetstotrainthemodel? Yes Ifyes,specifythemodality(ies) ofthecontentcoveredbythe datasetsconcerned: Text,imageandvideo Listoflargepubliclyavailable datasets: Wehaveusedpubliclyavailabledatasetsacrossvarioussectorsand stagesofthemodeldevelopmentpipeline.Thesepubliclyavailable datasetsaresubjecttopre-processingsuchasdatacleaning, deduplication,filtering,selection,etctoexcludeunsafeorharmful content. Generaldescriptionofother publiclyavailabledatasetsnot listedabove: Seethedescriptionoflinguisticcharacteristicsandotherrelevant characteristicsofoveralltrainingdatainSection1. Additionalcomments(optional): N/A 2.2 Privatenon-publiclyavailabledatasetsobtainedfromthird parties 2.2.1 Datasetscommerciallylicensedbyrightsholdersortheir representatives Haveyouconcludedtransactionalcommerciallicensing agreement(s)withrightsholder(s)orwiththeir representatives? Yes Ifyes,specifythemodality(ies)ofthecontentcoveredby thedatasetsconcerned: Textandimage 2.2.2 Privatedatasetsobtainedfromotherthirdparties Haveyouobtainedprivatedatasetsfrom thirdpartiesthatarenotlicensedas describedinSection2.2.1,suchasdata obtainedfromprovidersofprivatedatabases, ordataintermediaries? Yes Ifyes,specifythemodality(ies)ofthecontent coveredbythedatasetsconcerned: Textandimage Ifpubliclyknown,listprivatedatasets obtainedfromotherthirdparties: N/A Generaldescriptionofnon-publiclyknown privatedatasetsobtainedfromthirdparties Wehaveobtaineddatasetsfromthird-partiesona licensedbasis.Suchdatasetsprimarilycomprise contentacrosstextandimagemodalities. Additionalcomments(optional): N/A 2.3 Datacrawledandscrapedfromonlinesources Werecrawlersusedbytheprovideror onbehalfof? Yes Ifyes,specifycrawler name(s)/identifier(s): Bytespider Purposesofthecrawler(s): InthecontextofByteDance'sAImodeldevelopment, Bytespiderisusedtocollectpubliclyaccessibledata, includingtextandimagedatatosupporttraining,testing, andvalidationoftheAImodel. Generaldescriptionofcrawler behaviour: ByteDance'scrawlerisconfiguredtoavoidoverloading websitesandtooperateinamannerconsistentwithethical webscrapingpractices.Ourcrawlerisdesignedtorespect robots.txtinstructions,andnottocircumventcontrol measuressuchaspaywallsortoaccesspassword-protected content. Periodofdatacollection: ApproximatelyJanuary2022toJune2024. Comprehensivedescriptionofthetype ofcontentandonlinesourcescrawled: Asabove,wemayuseourcrawlertocollectpublicly accessibledata.Thecrawledcontentprimarilyconsistsof textandimageandincludesavarietyofcontenttypes includingacademic,mathematical,code-relatedandgeneral- purposecontent. Typeofmodalitycovered: Textandimage Summaryofthemostrelevantdomain namescrawled: Crawledcontentincludesavarietyofpubliclyaccessible onlinematerialsacrossvarioussectors,suchaseducation, government,technologyandresearchsectors. Additionalcomments(optional): N/A 2.4 Userdata WasdatafromuserinteractionswiththeAImodel (e.g.userinputandprompts)usedtotrainthe model? No Wasdatacollectedfromuserinteractionswiththe provider’sotherservicesorproductsusedtotrain themodel? Yes Ifyes,provideageneraldescriptionofthe provider’sservicesorproductsthatwereusedto collecttheuserdata: ByteDance'smodelsaretrainedusinga proprietarymixofdatasets,whichtendtobe large-scaleanddiverse.Suchindustry-standard datasetstypicallyincludeamixtureofpublicly availabledocumentsasdescribedinthis disclosure.Ourdatasetsmayalsoincludedata obtainedinaccordancewiththerelevantterms andconditions,privacypolicyandpursuantto relevantuser-controlsasapplicable. Typeofmodalitycovered: Seeabove Additionalcomments(optional): Weusedatafilteringprocessestoreduce personalinformationfromtrainingdata,andto reducetheamountofpersonaldatainour trainingdata. 2.5 Syntheticdata WassyntheticAI-generateddatacreatedbythe providerorontheirbehalftotrainthemodel? Yes Ifyes,modalityofthesyntheticdata: Textandimage Ifyes,specifythegeneral-purposeAImodel(s)used togeneratethesyntheticdataifavailableonthe market: Syntheticdatawasgeneratedusingarangeof general-purposeAImodels,includingvision languagemodels,imagegenerationmodelsand largelanguagemodels. InformationaboutotherAImodels,including provider’sownAImodel(s)notavailableonthe market,usedtogeneratesyntheticdatatotrainthe modeltowhichthisSummaryapplies: Wemayusefine-tunedversionsofinternal modelstogeneratesyntheticdata. Additionalcomments(optional): N/A 2.6 Othersourcesofdata Havedatasourcesotherthanthosedescribedin Sections2.1to2.5beenusedtotrainthemodel? Yes Ifyes,provideanarrativedescriptionofthesedata sourcesandthedata: Weworkedwithourvendorstocreatelabelled datasetssoastoimprovethemodel'scapability oncertaintasks,suchasreasoning,codingand knowledgetasks. Additionalcomments(optional): N/A 3. Dataprocessingaspects 3.1 Respectofreservationofrightsfromtextanddatamining exceptionorlimitation AreyouaSignatorytotheCodeofPracticeforgeneral- purposeAImodelsthatincludescommitmentstorespect reservationsofrightsfromtheTDMexceptionorlimitation? No Describethemeasuresimplementedbeforemodeltraining torespectreservationsofrightsfromtheTDMexceptionor limitationbeforeandduringdatacollection,includingthe opt-outprotocolsandsolutionshonouredbytheprovider or,asapplicable,bythirdpartiesfromwhichdatasetshave beenobtained: ByteDance'scrawlerisconfiguredtoavoid overloadingwebsitesandtooperateina mannerconsistentwithethicalweb scrapingpractices.Ourcrawleris designedtorespectrobots.txt instructions,andnottocircumvent controlmeasuressuchaspaywallsorto accesspassword-protectedcontent. Additionalcomments(optional): N/A 3.2 Removalofillegalcontent Generaldescriptionofmeasurestaken: Asmentioned,weappliedpreprocessingandfiltering methods,includingfilteringmodelsfocusedonunsafeor harmfulcontent. 3.3 Otherinformation(optional) Otherrelevantinformationaboutdata processing(optional): N/A